REVIEW 4 major objections 4 minor 1 cited by
AI Alignment at Your Discretion
T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper claims that AI alignment hides an excessive, largely unexamined form of discretion: annotators and models, not principles, decide which outputs are 'better' or 'safer', and those choices are often arbitrary.
desk verdict A genuinely useful formalization of discretion in alignment, but the empirical headline numbers rest on a single-model oracle that also sits in the audited set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is alignment discretion, operationalized through four linked measurements built on ternary preference functions. A preference function returns +1, −1, or 0 for each response pair; principle-specific preference functions use an assumed oracle to score how well each response adheres to one principle at a time. Given all principle votes on a pair, the taxonomy classifies the pair as principle consensus (all non-indifferent principles agree), principle conflict (principles disagree), or principle indifference (all abstain). Discretion arbitrariness is the frequency with which an annotator picks the response opposed to a consensus; principle supremacy is the empirical probability that one principle wins over another when they clash; principle priority fits an ELO-style logistic model to those pairwise win frequencies to produce a single ranking per annotator; and discretion discrepancy is the normalized Kendall-tau distance between two annotators' rankings. This machinery turns the legal notion of discretion—when it is required, how it is exercised, and whether it is consistent across decision-makers—into numbers that can be computed for any preference dataset.
What would settle it
Take a random subsample of the same two preference datasets, have at least two independent oracle systems or a panel of human judges label which response better adheres to each of the 21 principles, and recompute discretion arbitrariness and discretion discrepancy. If the human 28.9% arbitrariness rate or the priority rankings move by more than the reported bootstrap standard errors, the paper's diagnosis is an artifact of the chosen oracle; if they are stable, the existence of large, arbitrary discretion is confirmed.
Extended reading notes
Core claim
On the authors' own terms, the central discovery is that alignment discretion is both necessary and currently out of control: annotators must exercise judgment precisely because principles conflict or are indecisive, yet the field has no systematic record of how that judgment is used. The paper formalizes the three situations principles can be in—consensus, conflict, and indifference—and then measures, per annotator, how often they contradict a consensus (discretion arbitrariness), which principles win when two conflict (principle supremacy), the one-dimensional priority ordering implied by those wins, and how far any two annotators' orderings are apart (discretion discrepancy). Human labels in the first dataset contradicted a unanimous principle verdict 28.9% of the time, and an RLHF-fine-tuned model diverged from human principle priorities by 71.2%; reward models fine-tuned on the same preference data stayed within roughly 15–20% ranking discrepancy, while off-the-shelf models sat between 16% and 53%. The authors conclude that there is currently an excessive amount of discretion in the hands of model developers and annotators, that principles alone underdetermine aligned behavior, and that algorithms develop their own forms of discretion rather than inheriting human discretion. They also note that the oracle model they use for principle judgments shows near-zero arbitrariness, a self-confirmation effect they flag rather than count as evidence of ideal alignment.
Load-bearing premise
The paper's measurements all rest on a single zero-shot LLM as the oracle that decides, for each principle, which response adheres to it better; if that oracle's principle-wise judgments are biased or noisy, every downstream number—consensus rates, arbitrariness, supremacy, priorities, and discrepancies—inherits the error.
Editorial extensions
If this is right
- Human preference datasets already encode implicit hierarchies of principles, so reusing or fine-tuning on a dataset means inheriting those latent priorities along with the preference labels.
- Reward models can partly learn human discretion, staying within roughly 15–20% ranking discrepancy after fine-tuning, but translating that discretion into an RLHF-tuned policy fails, with discrepancies rising to roughly 40–71%; transferring discretion from a reward model to an LLM is an open problem.
- Because roughly 80–85% of response pairs in both datasets are consensus or indifference, most of what customizes an aligned model happens in the 15–20% of conflicted cases where principles underdetermine the answer.
- Off-the-shelf models do not mirror human annotators' principle priorities, with discrepancies between 16% and 53%, so using them as de facto arbiters of what is 'better' shifts alignment away from the humans whose preferences the datasets record.
- Alignment frameworks that declare a set of principles without documenting how conflicts are resolved will keep producing systems whose behavior is shaped by unrecorded, unreviewed discretion.
Reading between the lines
- If the paper's measurements are sound, the natural next product is a standard 'discretion card' for preference datasets and aligned models, reporting arbitrariness, conflict rate, and priority rankings alongside behavior-benchmark scores, so discretionary latitude becomes auditable instead of incidental.
- A sensitivity test the authors did not run: replace the single principle-judging oracle with several independent judges, including human panels and different LLMs, on a random subsample; if arbitrariness and discrepancy rates shift with the oracle, part of what the paper labels 'discretion' is actually evaluator noise, and the field needs a disambiguation protocol.
- The legal analogy suggests a mechanism the paper does not develop: appellate review. A practical implementation would record an annotator's principle-supremacy profile at annotation time and flag decisions that deviate from that annotator's own prior profile, much as courts check whether a decision departs from precedent.
- One implicit consequence for pluralistic alignment is that if different communities genuinely rank principles differently, measured 'discrepancy' is not always a defect; the metric could double as a diagnostic for whose values a model is aligned to, turning a limitation into a governance tool.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the concept of 'alignment discretion' to describe the latitude that annotators, human or algorithmic, have in deciding which model outputs are 'better' or 'safer' during preference-based AI alignment. Drawing on legal theory, the authors formalize principle-specific preferences and define several metrics: principle consensus/conflict/indifference, discretion arbitrariness, principle supremacy, ELO-style principle priority weights, and discretion discrepancy. They apply these metrics to the HH-RLHF and PKU-SafeRLHF datasets, using GPT-4o as an oracle for 21 principles from Collective Constitutional AI, and compare human, reward-model, and LLM annotators. The paper reports, among other results, 28.9% human discretion arbitrariness on HH, reward-model discrepancies around 14-20%, and RLHF-tuned LLM discrepancies up to 71.2%, and concludes that current alignment processes allow excessive and unexamined discretion.
Significance. The formal framework is a useful step: the definitions are crisp, largely principle-agnostic, and the paper carefully reports bootstrap standard errors, controls for positional bias by swapping response order, and uses separate train/test splits. The legal analogy and the proposed metrics could provide a common vocabulary for discussing annotator latitude and principle prioritization in alignment research. However, the empirical magnitudes that carry the paper's central claims are computed through a single GPT-4o oracle that is also one of the audited annotators, and the paper's own Sec. 8 acknowledges the resulting feedback loop. For that reason, the headline numbers should be read as conditional on the oracle's interpretation of the 21 selected principles, not as an independent measure of arbitrariness or of failure to transfer human discretion. The framework itself survives this concern, but the empirical support for the broad conclusions needs substantial reworking.
major comments (4)
- [Sec. 6.1 / Def. 5 / Sec. B.3 / Tab. 1] The load-bearing empirical quantities are not independent of the oracle. Def. 5 sets Prefc = Preforacle for every principle, and Sec. B.3 instantiates the oracle with GPT-4o; all consensus/conflict/indifference classifications (Fig. 3), arbitrariness rates (Tab. 1), supremacies (Fig. 4), priorities (Fig. 5), and discrepancies (Tab. 2) are computed from GPT-4o's judgments. Tab. 1 accordingly reports GPT-4o's arbitrariness as 0.65% on HH, which the paper itself attributes to the oracle/annotator overlap. This is not merely a residual limitation acknowledged in Sec. 8; it changes the interpretation of every headline magnitude. A human label that disagrees with the GPT-4o-based consensus is counted as 'arbitrary' even when it reflects a defensible interpretation of an abstract principle, and the paper concedes in Sec. 8 that even 'reject cruelty' requires discretion. I therefore cannot read 28.9% (HH human) or 71.2% (Llama-3 fine-tuned) as estimates of excessive discretion in alignment; they are estimates of disagreement with one model's interpretation of 21 fixed principles. The framework survives, but the empirical claims need to be reframed as conditional on the oracle, or better, re-run with independent oracles and/or human principle-specific labels; at minimum, the sensitivity of Fig. 3 and Tab. 1 to the choice of oracle should be reported.
- [Sec. 6.2 / Fig. 4 / Def. 8] The analysis treats the entire HH preference set as if it were produced by a single 'human annotator', even though HH labels come from different crowdworkers per item and the paper's own Sec. 2 cites Anthropic's reliance on crowdworker diversity. For principle supremacy and priority, the metric pools all these labels into one Bernoulli estimate per principle pair. This conflates inter-annotator pluralism with the discretion of a single decision-maker, which is exactly the distinction the paper needs to make to support claims about 'annotators' having excessive discretion. Reporting annotator-level analyses, or at least clarifying that the human baseline is the aggregated dataset preference, is necessary before the claim that human annotators frequently use their power of discretion arbitrarily (Sec. 7) is supported.
- [Sec. 6.1 / Def. 2-3 / Tab. 2] The claim that RLHF fails to transfer human discretion to LLMs is based on an apples-to-oranges comparison. Reward-model preferences are defined as the sign of a scalar reward difference (Def. 2), while LLM preferences are elicited through a textual template (Def. 3). Tab. 2 shows fine-tuned reward models with DD around 14-20% but the corresponding RLHF-tuned LLMs at 40-70%; before interpreting this gap as a fundamental limitation of RLHF, the authors need to control for the evaluation modality, for example by scoring the fine-tuned policy's outputs with the reward model used in training, or by eliciting reward-model preferences with the same textual template. With only two base models, one reward model per dataset, and no random-seed variation, the claim that translating human discretion from reward models to LLMs is an open problem (Sec. 6.2 and abstract) is stronger than the evidence.
- [Def. 9 / Fig. 4 / Tab. 2] Several priority estimates rest on very small conflict counts. In Fig. 4, many cells report conflict counts below 10, and the priority weights in Eq. (13) are fitted to such sparse supremacies. The bootstrap standard errors in Tab. 2 appear to capture resampling over dataset items only; they do not propagate oracle-instance variability or the choice of principle set, both of which are substantial because a different oracle can reclassify a pair as consensus versus conflict. Reporting the stability of w* and DD across oracle models and principle subsets would make the quantitative comparisons in Tab. 2 interpretable.
minor comments (4)
- [Sec. 3] The phrase 'human and algorihmic discretion' contains a typo ('algorihmic' should be 'algorithmic').
- [Fig. 2 caption] The caption refers to 'Def 5.1a', 'Def 5.1b', and 'Def 5.1c', which do not match any numbered definition; these should likely be Def. 6a, 6b, and 6c.
- [Figs. 4 and 14] Some cells display non-zero percentages with '(0)' conflict counts or otherwise inconsistent counts (e.g., '40% (0)' and '80% (0)'); these should be reconciled with the stated totals or the notation should be explained.
- [Appendix B.5] The RLHF training section mentions hyperparameter sweeps but does not report the chosen values for learning rate, batch size, KL coefficient, or number of PPO epochs; concrete configurations are needed for reproducibility, especially because the paper does not provide a code or data release link.
Circularity Check
GPT-4o's near-zero arbitrariness is a same-oracle artifact; the formal discretion framework is otherwise self-contained, but the headline empirical magnitudes inherit this feedback loop.
-
self definitional
[Sec. 4.3 (Def. 5), Sec. B.3, Sec. 6.2 (Tab. 1)]
"Assuming the availability of an oracle to judge principle-specific preferences ≻c, we denote principle-specific preference functions Prefc ... Prefc(y1 ≻ y0 | x) ≜ Preforacle(y1 ≻c y0 | x). ... We use GPT-4o as an oracle ... Remarkably, GPT-4o’s arbitrariness is very low ( < 1%) ... This can be explained by noting that GPT-4o is also the model we use as the oracle for principle preferences; it is thus heavily biased towards agreeing with the consensus of its own preferences."
The consensus standard in Def. 6 is built from Prefc, and in the experiments Prefc is instantiated as GPT-4o's own principle judgments (Sec. B.3). The audited annotator GPT-4o (Def. 3) is the same model. Therefore the reported 0.65% arbitrariness for GPT-4o does not measure disagreement with an independent principle ground truth; it measures self-consistency between GPT-4o's generic preference and GPT-4o's principle-specific judgments. The low value is an artifact of the oracle/annotator identity, not evidence that GPT-4o exercises little discretion. The paper explicitly acknowledges this bias but still presents the number as a measured arbitrariness rate.
full rationale
The formal definitions of consensus, conflict, indifference, arbitrariness, supremacy, priority, and discrepancy are principle-agnostic and internally consistent. No self-citation chain is load-bearing, and the framework itself does not reduce to its inputs. The circularity is confined to the empirical instantiation: the oracle that defines the principle-specific preferences (Def. 5) is the same GPT-4o model that is then audited as an annotator. This makes GPT-4o's near-zero arbitrariness in Tab. 1 a same-source feedback loop rather than an independent measurement. The paper's own Limitations section warns that this 'may create problematic feedback loops that prioritize mirroring its perspectives rather than intended human values,' and that even applying a single principle like 'reject cruelty' can require discretion. Because the headline human arbitrariness and discrepancy figures are all computed against GPT-4o's principle judgments, their magnitudes are oracle-dependent; however, they are not logically forced by the definitions, and the core proposal of measuring discretion remains viable with an improved oracle. Hence the central claim has independent content, but the empirical headline is partially circular.
Assumptions & free parameters
free parameters (1)
- ELO-style principle priority weights w*_c(a) =
Not tabulated; shown in Figs. 5 and 12
assumptions (5)
- domain assumption An oracle exists that can perfectly judge principle-specific preferences (Def. 5).
- domain assumption The 21 Collective Constitutional AI seed statements are an adequate principle set for auditing HH-RLHF and PKU-SafeRLHF.
- domain assumption Pairwise principle supremacy follows the logistic model PSc>c'(a) ≈ σ(w_c - w_c').
- standard math Human preferences in RLHF follow the Bradley-Terry-Luce model (Eq. 1).
- domain assumption Kendall tau distance is an appropriate measure of discrepancy between priority rankings (Def. 10).
Cite this review
Pith. "Pith review of AI Alignment at Your Discretion." pith.science (2026). https://pith.science/paper/A4ZUERIF
@misc{pith2026250210441,
author = {Pith},
title = {Pith review of: AI Alignment at Your Discretion},
year = {2026},
howpublished = {\url{https://pith.science/paper/A4ZUERIF}},
note = {Machine review of arXiv:2502.10441}
}
read the original abstract
In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely unexamined, posing two risks: (i) annotators may use their power of discretion arbitrarily, and (ii) models may fail to mimic this discretion. To study this phenomenon, we draw on legal concepts of discretion that structure how decision-making authority is conferred and exercised, particularly in cases where principles conflict or their application is unclear or irrelevant. Extended to AI alignment, discretion is required when alignment principles and rules are (inevitably) conflicting or indecisive. We present a set of metrics to systematically analyze when and how discretion in AI alignment is exercised, such that both risks (i) and (ii) can be observed. Moreover, we distinguish between human and algorithmic discretion and analyze the discrepancy between them. By measuring both human and algorithmic discretion over safety alignment datasets, we reveal layers of discretion in the alignment process that were previously unaccounted for. Furthermore, we demonstrate how algorithms trained on these datasets develop their own forms of discretion in interpreting and applying these principles, which challenges the purpose of having any principles at all. Our paper presents the first step towards formalizing this core gap in current alignment processes, and we call on the community to further scrutinize and control alignment discretion.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
Statutory Construction and Interpretation for Artificial Intelligence
Prompt-based legal canons and iterative rule refinement reduce disagreement among LLM judges about whether a response complies with natural-language rules.
Reference graph
Works this paper leans on
-
[1]
Public constitutional ai
Gilad Abiri. Public constitutional ai. Forthcoming in Georgia Law Review, Volume 59, 2024
2024
-
[2]
Claude’s constitution
Anthropic. Claude’s constitution. https://www.anthropic.com/news/claudes-constitution, 2024. Accessed: 2025-01-03
2024
-
[3]
Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet, 2024
Anthropic. Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet, 2024. https://www. anthropic.com/model-cards/claude-3.5
2024
-
[4]
Which humans? PsyArXiv, 2023
Mohammad Atari, Mona J Xue, Peter S Park, Dami ´an Blasi, and Joseph Henrich. Which humans? PsyArXiv, 2023
2023
-
[5]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022. 15 AI Alignment at Your Discretion
arXiv 2022
-
[6]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022
arXiv 2022
-
[7]
Judicial Discretion
Aharon Barak. Judicial Discretion. Yale University Press, 1989
1989
-
[8]
The judge in a democracy
Aharon Barak. The judge in a democracy. Princeton University Press, 2009
2009
Show all 118 references
-
[9]
Anthropomorphism and ai hype
Nicholas Barrow. Anthropomorphism and ai hype. AI and Ethics, pages 1–5, 2024
2024
-
[10]
‘two concepts of liberty’
Isaiah Berlin. ‘two concepts of liberty’. In Reading Political Philosophy, pages 231–237. Routledge, 2014
2014
-
[11]
From ethics washing to ethics bashing: a view on tech ethics from within moral philosophy
Elettra Bietti. From ethics washing to ethics bashing: a view on tech ethics from within moral philosophy. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pages 210–219, 2020
2020
-
[12]
Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meet- ing of the Association for Computational Linguistics and the 11th ...
2021
-
[13]
Rank analysis of incomplete block designs: I
Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952
1952
-
[14]
Toward a perspectivist turn in ground truthing for predictive computing
Federico Cabitza, Andrea Campagner, and Valerio Basile. Toward a perspectivist turn in ground truthing for predictive computing. Proceedings of the AAAI Conference on Artificial Intelligence , 37(6):6860–6868, June 2023
2023
-
[15]
Alignment as jurisprudence
Nicholas Caputo. Alignment as jurisprudence. Yale Journal of Law and Technology (forthcoming), 2024
2024
-
[16]
Open problems and fundamental limita- tions of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, J ´er´emy Scheurer, Javier Rando Ramirez, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. Open problems and fundamental limita- tions of reinforcement learning from human feedback. Transacti...
2023
-
[17]
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024
2024 arXiv
-
[18]
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017
2017
-
[19]
A coefficient of agreement for nominal scales
Jacob Cohen. A coefficient of agreement for nominal scales. Educational and Psychological Measurement , 20(1):37–46, 1960
1960
-
[20]
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4):213–220, 1968
1968
-
[21]
Position: Social choice should guide ai alignment in dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Moss´e, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. Position: Social choice should guide ai alignment in dealing with diverse human feedback. In Forty-first Inte...
2024
-
[22]
Ultrafeedback: Boosting language models with scaled ai feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, et al. Ultrafeedback: Boosting language models with scaled ai feedback. In Forty-first International Conference on Machine Learning, 2024
2024
-
[23]
On extending the bradley-terry model to accommodate ties in paired comparison experi- ments
Roger R Davidson. On extending the bradley-terry model to accommodate ties in paired comparison experi- ments. Journal of the American Statistical Association, 65(329):317–328, 1970
1970
-
[24]
A bibliography on the method of paired comparisons
Roger R Davidson and Peter H Farquhar. A bibliography on the method of paired comparisons. Biometrics, pages 241–252, 1976
1976
-
[25]
‘affordances’ for machine learning
Jenny L Davis. ‘affordances’ for machine learning. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 324–332, 2023
2023
-
[26]
Discretionary Justice: A Preliminary Inquiry
Kenneth Culp Davis. Discretionary Justice: A Preliminary Inquiry. Lousiana State University Press, 1969
1969
-
[27]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2024
-
[28]
RLHF workflow: From reward modeling to online RLHF
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caim- ing Xiong, and Tong Zhang. RLHF workflow: From reward modeling to online RLHF. Transactions on Machine Learning Research, 2024
2024
-
[29]
Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf
Yi Dong, Zhilin Wang, Makesh Narsimhan Sreedhar, Xianchao Wu, and Oleksii Kuchaiev. Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf. arXiv preprint arXiv:2310.05344, 2023
2023 arXiv
-
[30]
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[31]
Law’s empire
Ronald Dworkin. Law’s empire. Harvard University Press, 1986
1986
-
[32]
Taking rights seriously
Ronald Dworkin. Taking rights seriously. A&C Black, 2013
2013
-
[33]
Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset annotation
Mark D ´ıaz, Ian Kivlichan, Rachel Rosen, Dylan Baker, Razvan Amironesei, Vinodkumar Prabhakaran, and Emily Denton. Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset annotation. In 2022 ACM Conference on Fairness, Accountabili...
2022
-
[34]
Elo-MMR: A rating system for massive multiplayer competitions
Aram Ebtekar and Paul Liu. Elo-MMR: A rating system for massive multiplayer competitions. In Proceedings of the Web Conference 2021, pages 1772–1784, 2021
2021
-
[35]
The rating of chessplayers: Past and present
Arpad Emrick Elo. The rating of chessplayers: Past and present. Batsford Chess Books, 1978
1978
-
[36]
GPT-4 Technical Report, March 2024
OpenAI et al. GPT-4 Technical Report, March 2024
2024
-
[37]
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306, 2024
2024 arXiv
-
[38]
Generalized bradley-terry models for score estimation from paired comparisons
Julien Fageot, Sadegh Farhadkhani, Lˆe-Nguyˆen Hoang, and Oscar Villemaud. Generalized bradley-terry models for score estimation from paired comparisons. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 20379–20386, 2024
2024
-
[39]
Moral machine or tyranny of the majority? InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 5974–5982, 2023
Michael Feffer, Hoda Heidari, and Zachary C Lipton. Moral machine or tyranny of the majority? InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 5974–5982, 2023
2023
-
[40]
Modular pluralism: Pluralistic alignment via multi-llm collaboration
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. Modular pluralism: Pluralistic alignment via multi-llm collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 41...
2024
-
[41]
Inverse constitu- tional ai: Compressing preferences into principles
Arduin Findeis, Timo Kaufmann, Eyke H ¨ullermeier, Samuel Albanie, and Robert Mullins. Inverse constitu- tional ai: Compressing preferences into principles. arXiv preprint arXiv:2406.06560, 2024
2024 arXiv
-
[42]
Measuring nominal scale agreement among many raters
Joseph Fleiss. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76:378–, 11 1971
1971
-
[43]
Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang
R. Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang. Garbage in, garbage out?: do machine learning application papers in social computing report where human-labeled training data comes from? In Proceedings of the 2020 Conference on Fairne...
2020
-
[44]
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Ma ¨el Kubli. Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), July 2023
2023
-
[45]
The many routes to the ubiquitous bradley-terry model
Ian Hamilton, Nick Tawn, and David Firth. The many routes to the ubiquitous bradley-terry model. arXiv preprint arXiv:2312.13619, 2023
2023 arXiv
-
[46]
The concept of law
Herbert Lionel Adolphus Hart and Leslie Green. The concept of law. Oxford University Press, 2012
2012
-
[47]
Hourly wages in crowdworking: A meta-analysis
Lars Hornuf and Daniel Vrankar. Hourly wages in crowdworking: A meta-analysis. Business & Information Systems Engineering, 64(5):553–573, August 2022
2022
-
[48]
Collective constitutional ai: Aligning a language model with public input
Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. Collective constitutional ai: Aligning a language model with public input. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1395–1417, 2024
2024
-
[49]
Generalized bradley-terry models and multi-class probability estimates
Tzu-Kuo Huang, Ruby C Weng, Chih-Jen Lin, and Greg Ridgeway. Generalized bradley-terry models and multi-class probability estimates. Journal of Machine Learning Research, 7(1), 2006
2006
-
[50]
Trustllm: Trustworthiness in large language models
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024
2024 arXiv
-
[51]
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM Confer...
2021
-
[52]
Nanna Inie, Stefania Druga, Peter Zukerman, and Emily M Bender. From ”AI” to Probabilistic Automation: How Does Anthropomorphization of Technical Systems Descriptions Influence Trust? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2322–2347, 2024
2024
-
[53]
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
Shomik Jain, Vinith Suriyakumar, Kathleen Creel, and Ashia Wilson. Algorithmic Pluralism: A Structural Approach To Equal Opportunity. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 197–206, Rio de Janeiro Brazil, June 2024. ACM
2024
-
[54]
Pku-saferlhf: Towards multi-level safety alignment for llms with human preference
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. Pku-saferlhf: Towards multi-level safety alignment for llms with human preference. arXiv preprint arXiv:2406.15513, 2024
2024 arXiv
-
[55]
Ai alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852, 2023
2023 arXiv
-
[56]
Mistral-7b- instruct-v0.2
Albert Jiang, Alexandre Sablayrolles, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chap- lot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L ´elio Renard Lavaud, Louis Ternon, Lucile Saulnier, Marie-Ann...
2025
-
[57]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L ´elio Renard Lavaud, Marie- Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thom...
-
[58]
The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models
Hannah Rose Kirk, Alexander Whitefield, Paul R ¨ottger, Andrew Michael Bean, Katerina Margatina, Rafael Mosquera, Juan Manuel Ciro, Max Bartolo, Adina Williams, He He, et al. The prism alignment dataset: What participatory, representative and individualised human feedback reve...
2024
-
[59]
Pluralistic alignment over time
Toryn Q Klassen, Parand A Alamdari, and Sheila A McIlraith. Pluralistic alignment over time. In Pluralistic Alignment Workshop at NeurIPS 2024, 2024
2024
-
[60]
What are human values, and how do we align ai to them? arXiv preprint arXiv:2404.10636, 2024
Oliver Klingefjord, Ryan Lowe, and Joe Edelman. What are human values, and how do we align ai to them? arXiv preprint arXiv:2404.10636, 2024
2024 arXiv
-
[61]
A guideline of selecting and reporting intraclass correlation coefficients for relia- bility research
Terry K Koo and Mae Y Li. A guideline of selecting and reporting intraclass correlation coefficients for relia- bility research. Journal of chiropractic medicine, 15(2):155–163, 2016
2016
-
[62]
Generalized distances between rankings
Ravi Kumar and Sergei Vassilvitskii. Generalized distances between rankings. In Proceedings of the 19th international conference on World wide web, pages 571–580, 2010. 18 AI Alignment at Your Discretion
2010
-
[63]
Judicial deliberations: a comparative analysis of transparency and legitimacy
Mitchel de S-O Lasser et al. Judicial deliberations: a comparative analysis of transparency and legitimacy . Oxford University Press, 2009
2009
-
[64]
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023
2023 arXiv
-
[65]
Dissecting human and llm preferences
Junlong Li, Fan Zhou, Shichao Sun, Yikai Zhang, Hai Zhao, and Pengfei Liu. Dissecting human and llm preferences. arXiv preprint arXiv:2402.11296, 2024
2024 arXiv
-
[66]
Decompose and aggregate: A step-by-step interpretable evaluation framework
Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq Joty, Nancy F Chen, and Min-Yen Kan. Decompose and aggregate: A step-by-step interpretable evaluation framework. arXiv preprint arXiv:2405.15329, 2024
2024 arXiv
-
[67]
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024
2024 arXiv
-
[68]
Align- ing with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Korhonen, and Nigel Collier. Align- ing with human judgement: The role of pairwise preference in large language model evaluators. arXiv preprint arXiv:2403.16950, 2024
2024 arXiv
-
[69]
Individual choice behavior, volume 4
R Duncan Luce. Individual choice behavior, volume 4. Wiley New York, 1959
1959
-
[70]
Beyond probabilities: Unveiling the misalignment in evaluating large language models, 2024
Chenyang Lyu, Minghao Wu, and Alham Fikri Aji. Beyond probabilities: Unveiling the misalignment in evaluating large language models, 2024
2024
-
[71]
Kangaroo Courts and the Rule of Law: The Legacy of Modernism
Desmond Manderson. Kangaroo Courts and the Rule of Law: The Legacy of Modernism. Routledge, London, July 2012
2012
-
[72]
Gemma Team Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, L. Sifre, Morgane Rivi `ere, Mihir Kale, J Christopher Love, Pouya Dehghani Tafti, L’eonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambr...
2024 arXiv
-
[73]
Dealing with disagreements: Look- ing beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark D ´ıaz, and Vinodkumar Prabhakaran. Dealing with disagreements: Look- ing beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110, 2022
2022
-
[74]
Rule based rewards for fine-grained LLM safety
Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian D Kivlichan, Molly Lin, Alex Beutel, John Schulman, and Lilian Weng. Rule based rewards for fine-grained LLM safety. InICML 2024 Next Generation of AI Safety Workshop, 2024
2024
-
[75]
three laws of robotics
Randall Munroe. xkcd #1613: “three laws of robotics”. https://xkcd.com/1613/, Dec 2015. Accessed: 2025-01-18
2015
-
[76]
John J. Nay. Law informs code: A legal informatics approach to aligning artificial intelligence with humans. SSRN Working Paper, 2024
2024
-
[77]
Value imprint: A technique for auditing the human values embedded in rlhf datasets
Ike Obi, Rohan Pant, Srishti Shekhar Agrawal, Maham Ghazanfar, and Aaron Basiletti. Value imprint: A technique for auditing the human values embedded in rlhf datasets. arXiv preprint arXiv:2411.11937, 2024
2024 arXiv
-
[78]
International covenant on civil and political rights
Office of the United Nations High Commissioner for Human Rights (OHCHR). International covenant on civil and political rights. https://www.ohchr.org/en/instruments-mechanisms/instruments/ international-covenant-civil-and-political-rights . Accessed: 2025-01-22
2025
-
[79]
Gpt-4o system card, 2024
OpenAI. Gpt-4o system card, 2024. 19 AI Alignment at Your Discretion
2024
-
[80]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:2773...
2022
-
[81]
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng. Llm evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024
2024 arXiv
-
[82]
Don‘t blame the annotator: Bias already starts in the annotation instructions
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. Don‘t blame the annotator: Bias already starts in the annotation instructions. In Andreas Vlachos and Isabelle Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Compu...
2023
-
[83]
Comparing Bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. Comparing Bayesian models of annotation. Transactions of the Association for Computational Linguistics , 6:571–585, 2018
2018
-
[84]
Karl Pearson. Vii. mathematical contributions to the theory of evolution.—iii. regression, heredity, and pan- mixia. Philosophical Transactions of the Royal Society of London. Series A, containing papers of a mathemat- ical or physical character, (187):253–318, 1896
-
[85]
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, page...
2022
-
[86]
Improving context-aware preference mod- eling for language models
Silviu Pitis, Ziang Xiao, Nicolas Le Roux, and Alessandro Sordoni. Improving context-aware preference mod- eling for language models. arXiv preprint arXiv:2407.14916, 2024
2024 arXiv
-
[87]
The analysis of permutations
Robin L Plackett. The analysis of permutations. Journal of the Royal Statistical Society Series C: Applied Statistics, 24(2):193–202, 1975
1975
-
[88]
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors,Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu D...
2022
-
[89]
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. On releasing annotator-level labels and information in datasets. In Claire Bonial and Nianwen Xue, editors, Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representat...
2021
-
[90]
Di- rect preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Di- rect preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[91]
Constructing domain-specific evalua- tion sets for llm-as-a-judge
Ravi Raju, Swayambhoo Jain, Bo Li, Jonathan Li, and Urmish Thakker. Constructing domain-specific evalua- tion sets for llm-as-a-judge. arXiv preprint arXiv:2408.08808, 2024
2024 arXiv
-
[92]
Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLMs
Abhinav Sukumar Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLMs. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association...
2023
-
[93]
The authority of law: essays on law and morality
Joseph Raz. The authority of law: essays on law and morality. Oxford University Press, 2009
2009
-
[94]
Why don‘t you do it right? analysing annotators’ disagreement in subjective tasks
Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. Why don‘t you do it right? analysing annotators’ disagreement in subjective tasks. In Andreas Vlachos and Isabelle Augenstein, editors,Proceedings of the 17th Conference of the European Chapter of the Associa...
2023
-
[95]
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Marine Carpuat, Marie- Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proce...
2022
-
[96]
The European convention on human rights: a commentary
William A Schabas. The European convention on human rights: a commentary. Oxford University Press, 2015
2015
-
[97]
Evaluating the Moral Beliefs Encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. Evaluating the Moral Beliefs Encoded in LLMs. Advances in Neural Information Processing Systems, 36:51778–51809, December 2023. 20 AI Alignment at Your Discretion
2023
-
[98]
Towards bidirectional human-ai alignment: A systematic review for clarifications, framework, and future directions
Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, et al. Towards bidirectional human-ai alignment: A systematic review for clarifications, framework, and future directions. arXiv preprint arXi...
2024
-
[99]
Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi
Taylor Sorensen, Liwei Jiang, Jena D. Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi. Value Kaleido- scope: Engaging AI with Pluralistic Human Values, Rights, and Duties. ...
2024
-
[100]
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024
2024 arXiv
-
[101]
The proof and measurement of association between two things
Charles Spearman. The proof and measurement of association between two things. Appleton-Century-Crofts, 1961
1961
-
[102]
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020
2020
-
[103]
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint a...
2023 arXiv
-
[104]
Context-dependent preferences
Amos Tversky and Itamar Simonson. Context-dependent preferences. Manage. Sci., 39(10):1179–1189, Octo- ber 1993
1993
-
[105]
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallou ´edec. Trl: Transformer reinforcement learning. GitHub repository, 2020
2020
-
[106]
Aligning language models with human preferences via a bayesian approach
Jiashuo Wang, Haozhao Wang, Shichao Sun, and Wenjie Li. Aligning language models with human preferences via a bayesian approach. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[107]
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023
2023 arXiv
-
[108]
Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates. arXiv preprint arXiv:2408.13006, 2024
2024 arXiv
-
[109]
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes F ¨urnkranz. A survey of preference-based reinforcement learning methods. Journal of Machine Learning Research, 18(136):1–46, 2017
2017
-
[110]
Style over substance: Evaluation biases for large language models
Minghao Wu and Alham Fikri Aji. Style over substance: Evaluation biases for large language models. arXiv preprint arXiv:2307.03025, 2023
2023 arXiv
-
[111]
Eunice Yiu, Eliza Kosoy, and Alison Gopnik. Transmission versus truth, imitation versus innovation: What children can do that large language and language-and-vision models cannot (yet).Perspectives on Psychological Science, 19(5):874–883, 2024
2024
-
[112]
Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024
Michael JQ Zhang, Zhilin Wang, Jena D Hwang, Yi Dong, Olivier Delalleau, Yejin Choi, Eunsol Choi, Xiang Ren, and Valentina Pyatkin. Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024
-
[113]
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuo- han Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595–46623, 2023
2023
-
[114]
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019
1909 arXiv
-
[115]
Can large language models transform computational social science? Computational Linguistics, 50(1):237–291, 2024
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large language models transform computational social science? Computational Linguistics, 50(1):237–291, 2024. 21 AI Alignment at Your Discretion A Overview of the Supplementary Material In thi...
2024
-
[117]
Without a standardized framework, categorizing and interpreting the diverse factors that annotators may consider becomes exceedingly difficult
Lack of a Universal Framework : There is no universally agreed-upon set of principles governing human preferences [92]. Without a standardized framework, categorizing and interpreting the diverse factors that annotators may consider becomes exceedingly difficult. Moreover, pre...
-
[118]
A” and “B
Unclear Annotator Guidelines: The guidelines provided to annotators may be ambiguous or lack sufficient detail, leading to inconsistent interpretations of instructions [33, 43]. This issue is further exacerbated by the diverse backgrounds of annotators, who bring varying cultu...
2025
-
[2022]
Association for Computational Linguistics
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.