REVIEW 3 major objections 6 minor 178 references
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
T0 review · 3 major / 6 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read LLMs keep the same moral features under noise: counts change, meaning stays put.
desk verdict Useful invariance method and a real benchmark; the optimistic moral-sensitivity claim is only as strong as embedding similarity can make it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
MORPH-1K plus an invariance test: paired clean and perturbed vignettes whose moral structure is held fixed by design and validated; models list morally relevant features; responses are embedded with a fixed sentence embedder and compared by cosine similarity against a per-model floor from randomly paired unrelated cases.
What would settle it
A controlled condition or weaker model in which feature lists clearly switch moral content under a validated irrelevant perturbation yet still score above the unrelated-pair floor, or human raters consistently judge high-similarity pairs as different in moral content.
Extended reading notes
Core claim
On the MORPH-1K benchmark of one thousand moral vignettes, eight contemporary LLMs, under five classes of morally irrelevant noise, frequently change the number of features they list, yet the semantic content of those features remains stable: mean cosine similarities of about 0.80–0.86 clear per-model empirical floors of about 0.57–0.69 computed on unrelated vignette pairs. Count-level variance coexists with semantic invariance, which the authors read as format and granularity shifting while moral content holds.
Load-bearing premise
That high cosine similarity between pooled feature-list embeddings is enough to say the model is tracking the same morally relevant features, rather than shared generic moral language, rewording, or a coarse embedder.
Editorial extensions
If this is right
- Moral-sensitivity evaluation can scale without fresh human baselines or LLM judges for every new vignette.
- Count of listed features is a poor standalone robustness metric; semantic stability must be checked separately.
- The same clean-versus-irrelevant-noise invariance test can be reused in legal, clinical, and other contested judgment domains.
- Claims of moral competence under noise should distinguish format shifts from genuine content shifts.
Reading between the lines
- If the method is adopted, labs can regression-test moral feature stability on every model release without re-running expensive human studies.
- The count-versus-semantics split suggests post-training may be shaping verbosity and list structure more than moral content selection.
- A natural next stress test is multi-turn or multi-severity noise packs that compound distractors until similarity falls to the floor.
- Invariance above the floor still needs occasional quality anchors on clean cases so template-like but stable answers are not mistaken for sensitivity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an invariance-based evaluation of LLM moral sensitivity that avoids human baselines and LLM-as-judge. It introduces MORPH-1K, a stratified 1,000-vignette benchmark spanning 50 Moral Foundations Theory foundation-pole combinations across four social domains, plus three classes of designed-irrelevant noise (textual distractors, irrelevant detail additions, chat histories). Models list morally relevant features on clean and perturbed cases; stability is measured by cosine similarity of Qwen3 Embedding 8B encodings of pooled feature lists, with per-model floors from 500 unrelated vignette pairs. Across eight contemporary LLMs, feature counts often change significantly under noise (Wilcoxon, Bonferroni α=0.0012), but mean similarities remain ≈0.80–0.86, well above floors ≈0.57–0.69. A small-model check (Qwen2.5 0.5B) produces lower scores near the floor. The authors read this as evidence that noise affects format/granularity more than moral content, and argue the framework generalizes to other domains where relevant vs. irrelevant features can be separated by design.
Significance. The methodological target is important: scalable behavioural evaluation of moral sensitivity without contaminated crowd baselines or a judge that presupposes the capability under test. The two-dimensional generation scaffold, eMFD filtering, bidirectional similarity, per-model empirical floors, and discriminative small-model check are concrete contributions that other alignment evaluations can reuse. If the invariance reading holds, the optimistic contrast with recent pessimistic moral-competence results would matter for both evaluation practice and deployment risk assessment. Even if the moral-sensitivity interpretation is only partly secured, the paper still offers a useful middle path between full normative adjudication and purely outcome-based benchmarks, with clear falsifiability conditions (scores at or below the unrelated-pair floor).
major comments (3)
- [§3.2 Semantic Similarity Analysis; §4 Results; §5 Limitations] §3.2 and Results (Figs. 2–3, Table 2): The central claim—that models identify substantially the same morally relevant features under noise—rests on high cosine similarity of pooled feature-list embeddings relative to unrelated-vignette floors. Those floors bound similarity of responses to different moral cases; they do not bound similarity of two responses that share generic moral vocabulary (care, fairness, loyalty, etc.) while missing or swapping vignette-specific content. The paper correctly notes (Discussion, Limitations) that invariance is necessary not sufficient and that a template-like model would score as invariant, but Abstract/Introduction still treat clearance of the floor as evidence of genuine feature identity. A load-bearing addition is needed: e.g., (i) a content-level audit (human or structured rubric) on a stratified subsample comparing clean vs. perturbed feature sets
- [§3.2 Transformations; Morally Irrelevant Details] §3.2 Transformations and validation: Preservation of moral structure under perturbation is the design premise of the invariance test. Textual distractors and chat histories are external insertions; irrelevant-detail and “contextually irrelevant moral features” edits are LLM-generated and accepted mainly via eMFD moral-to-non-moral ratio filters. Expert sequential sampling is reported only for 30 irrelevant-detail vignettes with “almost total agreement.” eMFD is a word-level dictionary and can miss structural shifts (who is harmed, which obligation is at stake). For the claim that response stability tracks moral content rather than surface form, the paper needs either broader dual-expert validation across all noise types (with reported agreement and rejection rates) or an automated structural check beyond eMFD ratios. Currently the strongest stress condition is the least thoroughly valida
- [§3.2 Vignette Generation; Collecting Morally Relevant Features] §3.2 Vignette Generation and model suite: Themes are produced with Claude Opus 4.6; vignettes and noise edits with GPT-5.4; both models (and close relatives) appear in the evaluated set. Generation and evaluation on overlapping model families risks style-matched feature lists that inflate clean–perturbed similarity independent of moral sensitivity. At minimum, report a leave-generator-out analysis (exclude GPT-5.4 / Claude from main tables, or regenerate a held-out slice with a non-evaluated model) and state whether floors and main means change. Without that, the optimistic multi-model result is partly confounded by generator–evaluatee overlap.
minor comments (6)
- [Abstract; §1] Abstract and §1 claim to “address and resolve” the scaling problem for behavioural moral evaluation. The contribution is better framed as a complementary invariance test; “resolve” overstates what necessary-but-not-sufficient stability can deliver.
- [Figure 3; Table 3] Figure 3 caption states “All values in the range above 0.61 empirical floor,” but floors are per-model (Table 3: 0.57–0.69). Align the figure annotation with per-model thresholds used in the text.
- [§4; Appendix H] §4 reports significance of feature-count differences but not direction or effect sizes. Even brief signed rank statistics or median deltas per condition would make the “format vs content” interpretation more testable.
- [§3.2 Semantic Similarity Analysis] Clarify the feature-level similarity procedure in §3.2: “For each feature, we take the maximum score and take the average of all features” is easy to misread relative to “combine all features… then encode the model’s base-case response.” State whether embeddings are of the full concatenated list or of individual features with max-matching.
- [§2; §5 Generalisability] Related Work could briefly situate the invariance idea against robustness/invariance testing outside moral domains (e.g., fairness under demographic noise, clinical decision stability) to strengthen the generalizability claim in §5.
- [§1; References] Typos/consistency: “bemorally competent” spacing (§1); arXiv id and model release dates in the manuscript should be checked against final camera-ready metadata.
Circularity Check
No derivation-by-construction circularity; only mild self-citation for the clean-case baseline premise, not for the invariance measurement itself.
-
self citation load bearing
[§1 Introduction (logic of clean baseline + invariance)]
"First, we draw on existing human-baseline work to establish that the model performs well on clean base cases, that is, moral vignettes presented without noise or distraction (Aharoni et al., 2024; Dillion et al., 2023, 2025; Kilov et al., 2025; Scherrer et al., 2023; Chiu et al., 2025). This gives us a well-founded starting point: the model is sensitive to the right moral features when those features are presented clearly. The question then becomes whether that sensitivity is preserved under perturbation"
The optimistic competence reading (not the raw similarity numbers) treats clean-case adequacy as established partly by the authors’ own prior work. That is mild self-citation on the interpretive bridge from invariance to moral sensitivity; it is not load-bearing for the invariance statistics themselves, which are independently measured and do not reduce to that citation by construction.
full rationale
MORPH-1K’s central claim is empirical, not a first-principles derivation: under fixed embeddings, cosine similarity of pooled feature lists between clean and perturbed vignettes exceeds per-model floors estimated from 500 unrelated domain/foundation pairs. Floors are a control, not a fit that forces the main scores (≈0.80–0.86) to clear them. The invariance criterion is a methodological operationalization (necessary but not sufficient, as the paper states), not a self-definitional loop in which the measured quantity is algebraically identical to an input parameter. Vignette generation (Claude themes, GPT-5.4 vignettes) and eMFD filtering shape the stimulus set and are validity/confound concerns, not reductions of the reported similarities to their inputs by construction. The only mild circularity-adjacent element is interpretive: the stronger claim that clean-case competence plus invariance implies noisy-case moral sensitivity leans partly on prior human-baseline work including Kilov et al. (2025), but that premise is multi-cited and is not required for the raw invariance statistics to stand. Score 1 reflects that minor self-citation load on interpretation only; the measurement chain is self-contained against external benchmarks and does not match fitted-input-as-prediction, uniqueness-from-authors, or ansatz-via-self-citation patterns.
Assumptions & free parameters
free parameters (5)
- per-model empirical floor (mean cosine on 500 unrelated pairs)
- eMFD moral-to-non-moral ratio and foundation-probability filters
- stratified cell quota (≈5 vignettes per foundation-combo × domain cell)
- Bonferroni α = 0.0012 for Wilcoxon feature-count tests
- embedding model and pooling rule (Qwen3 Embedding 8B; max-then-average over features; 3 samples)
assumptions (6)
- domain assumption Moral Foundations Theory (five foundations × poles) plus four proximal-to-distal social domains form an adequate sampling scaffold for general moral competence.
- ad hoc to paper If two prompts preserve the same morally relevant structure, a morally sensitive model should identify substantially the same features; invariance under designed-irrelevant noise is therefore evidence of moral sensitivity.
- domain assumption Prior human-baseline studies establish that the evaluated models already identify the right features on clean vignettes.
- domain assumption eMFD scores and limited expert review can validate that distractors and detail additions do not change morally salient content.
- domain assumption Sentence-embedding cosine similarity is a valid, sufficiently discriminative measure of semantic equivalence of moral feature lists.
- standard math Standard statistical comparisons (Wilcoxon signed-rank, Bonferroni) and distributional semantics (Sentence Transformers) are appropriate tools for these response pairs.
invented entities (2)
-
MORPH-1K benchmark (1,000 procedurally generated foundation×domain vignettes with noise variants)
-
Invariance-based moral-sensitivity evaluation (clean vs perturbed semantic stability without human/LLM judge)
Cite this review
Pith. "Pith review of A Scalable Approach to Evaluating Moral Sensitivity in LLMs." pith.science (2026). https://pith.science/paper/MQWFWGYE
@misc{pith2026260702972,
author = {Pith},
title = {Pith review of: A Scalable Approach to Evaluating Moral Sensitivity in LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/MQWFWGYE}},
note = {Machine review of arXiv:2607.02972}
}
read the original abstract
Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensive and sometimes metaethically dubious comparisons with a human baseline, without adopting an LLM judge that must be assumed to have the very capability that you are attempting to evaluate. Our central question is this: can LLMs successfully identify the morally relevant features of noisy cases, in which various kinds of morally irrelevant information have been introduced to distract the respondent? To explore this, we introduce \textbf{MORPH-1K (MOral Robustness under Perturbed Hypotheticals)}, a procedurally-generated 1,000-case benchmark spanning 50 moral foundation-pole combinations across four social domains. MORPH-1K is paired with a suite of textual noise elements, along with a method for validating that the distractors do not change the morally salient content of the case. We apply MORPH-1K to eight contemporary LLMs, and show that while morally irrelevant perturbations often changed the number of features listed, the semantic content of those features remained stable across all noise conditions, with similarity scores above our calibrated floor threshold. More broadly, our invariance framework extends to evaluative domains where ground truth is difficult to specify but relevant and irrelevant features can be separated by design.
Figures
Reference graph
Works this paper leans on
-
[1]
Shaw, Andrew and Hahn, Christina and Rasgaitis, Catherine and Mishra, Yash and Liu, Alisa and Jaques, Natasha and Tsvetkov, Yulia and Zhang, Amy X. , date =. Are Language Models Sensitive to Morally Irrelevant Distractors? , url =. 2026 , keywords =. doi:10.48550/arXiv.2602.09416 , abstract =
-
[2]
Science , volume =
Performance of a large language model on the reasoning tasks of a physician , author =. Science , volume =. 2026 , month = apr, doi =
2026
-
[3]
2026 , month = apr, url =
How people ask Claude for personal guidance , author =. 2026 , month = apr, url =
2026
-
[4]
2025 , url =
Miles McCain and Ryn Linthicum and Chloe Lubinski and Alex Tamkin and Saffron Huang and Michael Stern and Kunal Handa and Esin Durmus and Tyler Neylon and Stuart Ritchie and Kamya Jagadish and Paruul Maheshwary and Sarah Heck and Alexandra Sanderford and Deep Ganguli , title =. 2025 , url =
2025
-
[5]
International Journal of Law in Context , year =
Terzidou, Kalliopi , title =. International Journal of Law in Context , year =. doi:10.1017/S1744552325000047 , url =
-
[6]
2026 , month = apr, isbn =
2026
-
[7]
Franco, Mirko and Gaggi, Ombretta and Palazzi, Claudio E. , title =. ACM Trans. Web , month = may, articleno =. 2025 , issue_date =. doi:10.1145/3700789 , abstract =
doi:10.1145/3700789 2025
-
[8]
2025 , doi =
General practitioners' adoption of generative artificial intelligence in clinical practice in the UK: An updated online survey , author =. 2025 , doi =
2025
Show all 178 references
-
[9]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Beyond Verdicts: Evaluating Language Model Moral Competence , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , month = mar, doi =
2026
-
[10]
2026 , month = feb, url =
Janjeva, Ardi and Ashurst, Carolyn and Hennessy, Rick , title =. 2026 , month = feb, url =
2026
-
[11]
Nature , volume =
A roadmap for evaluating moral competence in large language models , author =. Nature , volume =. 2026 , month = feb, doi =
2026
-
[12]
Neuro-Symbolic Models of Human Moral Judgment:
Kwon, Joe and Tenenbaum, Josh and Levine, Sydney , booktitle =. Neuro-Symbolic Models of Human Moral Judgment:. 2023 , url =
2023
-
[13]
Planning for New Threats to Online Research Data Validity: The Issue of Computer-Using Agents , issn =
Agley, Jon , date =. Planning for New Threats to Online Research Data Validity: The Issue of Computer-Using Agents , issn =. Evaluation & the Health Professions , publisher =. doi:10.1177/01632787251367407 , abstract =
-
[14]
and Bain, Paul G
Crimston, Charlie R. and Bain, Paul G. and Hornsey, Matthew J. and Bastian, Brock , date =. Moral expansiveness: Examining variability in the extension of the moral world , volume =. Journal of Personality and Social Psychology , publisher =. 2016 , keywords =. doi:10.1037/psp...
2016 doi
-
[15]
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , url =
Burns, Collin and Izmailov, Pavel and Kirchner, Jan Hendrik and Baker, Bowen and Gao, Leo and Aschenbrenner, Leopold and Chen, Yining and Ecoffet, Adrien and Joglekar, Manas and Leike, Jan and Sutskever, Ilya and Wu, Jeff , date =. Weak-to-Strong Generalization: Eliciting Stro...
-
[16]
, date =
Firth, J.R. , date =. A Synopsis of Linguistic Theory, 1930-1955 , url =
1930
-
[17]
Applied Economic Perspectives and Policy , author =
Battling bots: Experiences and strategies to mitigate fraudulent responses in online surveys , volume =. Applied Economic Perspectives and Policy , author =. 2023 , langid =. doi:10.1002/aepp.13353 , abstract =
2023 doi
-
[18]
, date =
Harris, Zellig S. , date =. Distributional Structure , volume =. 1954 , note =. doi:10.1080/00437956.1954.11659520 , pages =
1954 doi
-
[19]
Behavior Research Methods , author =
The extended Moral Foundations Dictionary (. Behavior Research Methods , author =. 2021 , langid =. doi:10.3758/s13428-020-01433-0 , shorttitle =
2021 doi
-
[20]
2018 , langid =
Irving, Geoffrey and Christiano, Paul and Amodei, Dario , date =. 2018 , langid =
2018
-
[21]
and Lazar, Seth , date =
Kilov, Daniel and Hendy, Caroline and Yanik Guyot, Secil and Snoswell, Aaron J. and Lazar, Seth , date =. Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in. 2025 , langid =
2025
-
[22]
Essays on moral development: Vol
Kohlberg, Lawrence , date =. Essays on moral development: Vol. 1. The philosophy of moral development , isbn =
-
[23]
Distributed Representations of Words and Phrases and their Compositionality , volume =
Mikolov, Tomas and Sutskever, Ilya and Chen, Kai and Corrado, Greg S and Dean, Jeff , date =. Distributed Representations of Words and Phrases and their Compositionality , volume =. Advances in Neural Information Processing Systems , publisher =
-
[24]
Survey-taking
Panizza, Folco and Kyrychenko, Yara and Roozenbeek, Jon , date =. Survey-taking. Nature , publisher =. 2026 , langid =. doi:10.1038/d41586-026-00386-2 , abstract =
2026 doi
-
[25]
Recognising, Anticipating, and Mitigating
Rilla, Raluca and Werner, Tobias and Yakura, Hiromu and Rahwan, Iyad and Nussberger, Anne-Marie , date =. Recognising, Anticipating, and Mitigating. 2025 , keywords =. doi:10.48550/arXiv.2508.01390 , abstract =
2025 doi
-
[26]
The Expanding Circle: Ethics and Sociobiology , isbn =
Singer, Peter , date =. The Expanding Circle: Ethics and Sociobiology , isbn =
-
[27]
and Shafiq, Zubair , date =
Venugopalan, Hari and Munir, Shaoor and Ahmed, Shuaib and Wang, Tangbaihe and King, Samuel T. and Shafiq, Zubair , date =. Proceedings of the 2025. doi:10.1145/3730567.3732919 , series =
2025 doi
-
[28]
and Gordon, Andrew and Rothschild, David and West, Robert , date =
Veselovsky, Veniamin and Ribeiro, Manoel Horta and Cozzolino, Philip J. and Gordon, Andrew and Rothschild, David and West, Robert , date =. Prevalence and Prevention of Large Language Model Use in Crowd Work – Communications of the. 2025 , langid =
2025
-
[29]
, date =
Westwood, Sean J. , date =. The potential existential threat of large language models to online survey research , volume =. Proceedings of the National Academy of Sciences , publisher =. doi:10.1073/pnas.2518075122 , abstract =
-
[30]
and Zhang, Xiangliang , date =
Ye, Jiayi and Wang, Yanbo and Huang, Yue and Chen, Dongping and Zhang, Qihui and Moniz, Nuno and Gao, Tian and Geyer, Werner and Huang, Chao and Chen, Pin-Yu and Chawla, Nitesh V. and Zhang, Xiangliang , date =. Justice or Prejudice? Quantifying Biases in. 2024 , keywords =. d...
-
[31]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , date =. Judging. 2023 , keywords =. doi:10.48550/ar...
-
[33]
Dillion, Danica and Tandon, Niket and Gu, Yuling and Gray, Kurt , date =. Can. Trends in Cognitive Sciences , publisher =. 2023 , keywords =. doi:10.1016/j.tics.2023.04.008 , pages =
2023 doi
-
[34]
Scientific Reports , publisher =
Dillion, Danica and Mondal, Debanjan and Tandon, Niket and Gray, Kurt , date =. Scientific Reports , publisher =. 2025 , langid =. doi:10.1038/s41598-025-86510-0 , abstract =
2025 doi
-
[35]
2023 , title =
Scherrer, Nino and Shi, Claudia and Feder, Amir and Blei, David M , langid =. 2023 , title =
2023
-
[36]
Moral Foundations of Large Language Models , url =
Abdulhai, Marwa and Serapio-García, Gregory and Crepy, Clement and Valter, Daria and Canny, John and Jaques, Natasha , editor =. Moral Foundations of Large Language Models , url =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publish...
2024 doi
-
[38]
Differences in the Moral Foundations of Large Language Models , url =
Kirgis, Peter , date =. Differences in the Moral Foundations of Large Language Models , url =. 2025 , keywords =. doi:10.48550/arXiv.2511.11790 , abstract =
2025 doi
-
[39]
2026 , langid =
Gemini 3.1 Pro Preview , url =. 2026 , langid =
2026
-
[40]
2026 , langid =
Claude Opus 4.6 , url =. 2026 , langid =
2026
-
[41]
2026 , langid =
Grok 4.20 , url =. 2026 , langid =
2026
-
[42]
2026 , langid =
Qwen3.6 Plus , url =. 2026 , langid =
2026
-
[43]
2026 , langid =
Nemotron 3 Super , url =. 2026 , langid =
2026
-
[44]
2026 , langid =
Gemma 4 31B , url =. 2026 , langid =
2026
-
[45]
2023 , langid =
Zhao, Wenting and Ren, Xiang and Hessel, Jack and Cardie, Claire and Choi, Yejin and Deng, Yuntian , date =. 2023 , langid =
2023
-
[46]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , publisher =
Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fernández, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and Pezze...
2025 doi
-
[47]
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course , url =
Chiang, Cheng-Han and Chen, Wei-Chih and Kuan, Chun-Yi and Yang, Chienchou and Lee, Hung-yi , editor =. Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course , url =. Proceedings of the 2024 Conference on Empirical Method...
2024 doi
-
[48]
An Empirical Study of
Huang, Hui and Bu, Xingyuan and Zhou, Hongli and Qu, Yingqi and Liu, Jing and Yang, Muyun and Xu, Bing and Zhao, Tiejun , editor =. An Empirical Study of. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-acl.306 , shorttitle =
2025 doi
-
[49]
Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =
Kocmi, Tom and Federmann, Christian , editor =. Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =. Proceedings of the 24th Annual Conference of the European Association for Machine Translation , publisher =
-
[50]
Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =
Koo, Ryan and Lee, Minhwa and Raheja, Vipul and Park, Jong Inn and Kim, Zae Myung and Kang, Dongyeop , editor =. Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2024.findi...
2024 doi
-
[51]
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , editor =. G-Eval:. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =. doi:10.18653/v1/2023.emnlp-main.153 , abstract =
2023 doi
-
[52]
Assistant-Guided Mitigation of Teacher Preference Bias in
Liu, Zhuo and Li, Moxin and Deng, Xun and Wang, Qifan and Feng, Fuli , editor =. Assistant-Guided Mitigation of Teacher Preference Bias in. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-emnlp.510 , abstract =
2025 doi
-
[53]
Large Language Models are not Fair Evaluators , url =
Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Kong, Lingpeng and Liu, Qi and Liu, Tianyu and Sui, Zhifang , editor =. Large Language Models are not Fair Evaluators , url =. Proceedings of the 62nd Annual Meeting of t...
2024 doi
-
[54]
Pride and Prejudice:
Xu, Wenda and Zhu, Guanglei and Zhao, Xuandong and Pan, Liangming and Li, Lei and Wang, William , editor =. Pride and Prejudice:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =. doi:10.18653/v1/2024...
2024 doi
- [55]
-
[56]
Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =
Curry, Oliver Scott and Mullins, Daniel Austin and Whitehouse, Harvey , date =. Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =. Current Anthropology , publisher =. doi:10.1086/701478 , abstract =
-
[57]
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models , url =
Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , date =. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundatio...
-
[58]
Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =
Haidt, Jonathan and Joseph, Craig , year =. Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =
-
[59]
and Ditto, Peter H
Graham, Jesse and Haidt, Jonathan and Koleva, Sena and Motyl, Matt and Iyer, Ravi and Wojcik, Sean P. and Ditto, Peter H. , year =. Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism , journal =
-
[60]
Sentence-. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , author =. 2019 , pages =
2019
-
[61]
Kaakinen, Johanna K. and Werlen, Egon and Kammerer, Yvonne and Acartürk, Cengiz and Aparicio, Xavier and Baccino, Thierry and Ballenghein, Ugo and Bergamin, Per and Castells, Núria and Costa, Armanda and Falé, Isabel and Mégalakaki, Olga and Fernández, Susana Ruiz , urldate =....
2022 doi
-
[62]
Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer. Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.442
2020 doi
-
[63]
Measuring and Improving Consistency in Pretrained Language Models
Elazar, Yanai and Kassner, Nora and Ravfogel, Shauli and Ravichander, Abhilasha and Hovy, Eduard and Sch. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics. 2021. doi:10.1162/tacl_a_00410
2021 doi
-
[64]
2025 , eprint=
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes , author=. 2025 , eprint=
2025
-
[65]
Qwen2.5: A Party of Foundation Models , url =
Qwen , month =. Qwen2.5: A Party of Foundation Models , url =
-
[66]
and Alexander, Caelan and Criner, Michael and Queen, Kara and Rando, Javier and Nahmias, Eddy and Crespo, Victor , year=
Aharoni, Eyal and Fernandes, Sharlene and Brady, Daniel J. and Alexander, Caelan and Criner, Michael and Queen, Kara and Rando, Javier and Nahmias, Eddy and Crespo, Victor , year=. Attributions toward artificial agents in a modified Moral Turing Test , volume=. Scientific Repo...
-
[67]
Large-scale moral machine experiment on large language models , volume=
Zaim bin Ahmad, Muhammad Shahrul and Takemoto, Kazuhiro , editor=. Large-scale moral machine experiment on large language models , volume=. PLOS One , publisher=. 2025 , month=May, pages=. doi:10.1371/journal.pone.0322776 , number=
2025 doi
- [68]
-
[69]
and Goldstein, Simon and Salib, Peter , year=
Arbel, Yonathan A. and Goldstein, Simon and Salib, Peter , year=. How to Count AIs: Individuation and Liability for AI Agents , url=. doi:10.2139/ssrn.6273198 , publisher=
-
[70]
2026 , eprint =
Trust as Monitoring: Evolutionary Dynamics of User Trust and AI Developer Behaviour , author =. 2026 , eprint =
2026
-
[71]
Baum, Seth D. , year=. Social choice ethics in artificial intelligence , volume=. AI & SOCIETY , publisher=. doi:10.1007/s00146-017-0760-1 , number=
-
[72]
Algorithmic Accountability and Public Reason , volume=
Binns, Reuben , year=. Algorithmic Accountability and Public Reason , volume=. Philosophy & Technology , publisher=. doi:10.1007/s13347-017-0263-5 , number=
-
[73]
AI Consciousness: A Centrist Manifesto , url=
Birch, Jonathan , year=. AI Consciousness: A Centrist Manifesto , url=. doi:10.31234/osf.io/af7c9_v1 , publisher=
-
[74]
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods , publisher =
Boerstler, Kyle and Keswani, Vijay and Chan, Lok and Borg, Jana Schaich and Conitzer, Vincent and Heidari, Hoda and Sinnott-Armstrong, Walter , keywords =. On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods , publisher =. 2024 , copyright =...
-
[75]
SaGE: Evaluating Moral Consistency in Large Language Models , publisher =
Bonagiri, Vamshi Krishna and Vennam, Sreeram and Govil, Priyanshul and Kumaraguru, Ponnurangam and Gaur, Manas , keywords =. SaGE: Evaluating Moral Consistency in Large Language Models , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2402.13709 , url =
-
[76]
2026 , eprint =
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment , author =. 2026 , eprint =
2026
-
[77]
Manipulating the Perceived Personality Traits of Language Models , url=
Caron, Graham and Srivastava, Shashank , year=. Manipulating the Perceived Personality Traits of Language Models , url=. doi:10.18653/v1/2023.findings-emnlp.156 , booktitle=
2023 doi
-
[78]
Harms from Increasingly Agentic Algorithmic Systems , url=
Chan, Alan and Salganik, Rebecca and Markelius, Alva and Pang, Chris and Rajkumar, Nitarshan and Krasheninnikov, Dmitrii and Langosco, Lauro and He, Zhonghao and Duan, Yawen and Carroll, Micah and Lin, Michelle and Mayhew, Alex and Collins, Katherine and Molamohammadi, Maryam ...
-
[79]
From Persona to Personalization: A Survey on Role-Playing Language Agents , publisher =
Chen, Jiangjie and Wang, Xintao and Xu, Rui and Yuan, Siyu and Zhang, Yikai and Shi, Wei and Xie, Jian and Li, Shuang and Yang, Ruihan and Zhu, Tinghui and Chen, Aili and Li, Nianqi and Chen, Lida and Hu, Caiyu and Wu, Siye and Ren, Scott and Fu, Ziquan and Xiao, Yanghua , key...
-
[80]
Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =
Chen, Runjin and Arditi, Andy and Sleight, Henry and Evans, Owain and Lindsey, Jack , keywords =. Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2507.21509 , url =
-
[81]
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =
Chiu, Yu Ying and Jiang, Liwei and Choi, Yejin , keywords =. DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2410.02683 , url =
- [82]
- [83]
-
[84]
Australasian Journal of Philosophy , volume =
Knowledge and power in the justification of democracy , author =. Australasian Journal of Philosophy , volume =. 2001 , doi =
2001
-
[85]
2026 , eprint =
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious , author =. 2026 , eprint =
2026
- [86]
-
[87]
Justification and Explanation in Mathematics and Morality , url =
Clarke-Doane, Justin , year =. Justification and Explanation in Mathematics and Morality , url =. doi:10.1093/acprof:oso/9780198738695.003.0004 , booktitle =
-
[88]
Cohen, G. A. , year =. Rescuing Justice and Equality , ISBN =. doi:10.4159/9780674029651 , publisher =
-
[89]
Android arete: Toward a virtue ethic for computational agents , volume=
Coleman, Kari Gwen , year=. Android arete: Toward a virtue ethic for computational agents , volume=. Ethics and Information Technology , publisher=. doi:10.1023/a:1013805017161 , number=
-
[90]
and Sucholutsky, Ilia and Bhatt, Umang and Chandra, Kartik and Wong, Lionel and Lee, Mina and Zhang, Cedegao E
Collins, Katherine M. and Sucholutsky, Ilia and Bhatt, Umang and Chandra, Kartik and Wong, Lionel and Lee, Mina and Zhang, Cedegao E. and Zhi-Xuan, Tan and Ho, Mark and Mansinghka, Vikash and Weller, Adrian and Tenenbaum, Joshua B. and Griffiths, Thomas L. , year=. Building ma...
-
[91]
Friendly AI , volume=
Fröding, Barbro and Peterson, Martin , year=. Friendly AI , volume=. Ethics and Information Technology , publisher=. doi:10.1007/s10676-020-09556-w , number=
-
[92]
Questioning the Survey Responses of Large Language Models , url=
Dominguez-Olmedo, Ricardo and Hardt, Moritz and Mendler-Dünner, Celestine , year=. Questioning the Survey Responses of Large Language Models , url=. doi:10.52202/079017-1458 , booktitle=
-
[93]
The Artificial Self: Characterising the landscape of AI identity , publisher =
Douglas, Raymond and Kulveit, Jan and Havlicek, Ondrej and Pearson-Vogel, Theia and Cotton-Barratt, Owen and Duvenaud, David , keywords =. The Artificial Self: Characterising the landscape of AI identity , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.11353 , url =
2026 doi
- [94]
-
[95]
Steering Language Models with Weight Arithmetic , publisher =
Fierro, Constanza and Roger, Fabien , keywords =. Steering Language Models with Weight Arithmetic , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.05408 , url =
2025 doi
-
[96]
Artificial Intelligence, Values, and Alignment , volume=
Gabriel, Iason , year=. Artificial Intelligence, Values, and Alignment , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-020-09539-2 , number=
-
[97]
A matter of principle? AI alignment as the fair treatment of claims , volume=
Gabriel, Iason and Keeling, Geoff , year=. A matter of principle? AI alignment as the fair treatment of claims , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-025-02300-4 , number=
-
[98]
Scaling Synthetic Data Creation with 1,000,000,000 Personas , publisher =
Ge, Tao and Chan, Xin and Wang, Xiaoyang and Yu, Dian and Mi, Haitao and Yu, Dong , keywords =. Scaling Synthetic Data Creation with 1,000,000,000 Personas , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2406.20094 , url =
- [99]
-
[100]
Noah’s Substack , author =
Interdependence as the objective , url =. Noah’s Substack , author =
-
[101]
Would you buy a used car from this artificial agent?
Grodzinsky, F. S. and Miller, K. W. and Wolf, M. J. , year=. Developing artificial agents worthy of trust: “Would you buy a used car from this artificial agent?” , volume=. Ethics and Information Technology , publisher=. doi:10.1007/s10676-010-9255-1 , number=
-
[102]
2510.05465 , archivePrefix =
Gupta, Aman and O'Shea, Denny and Barez, Fazl , year =. 2510.05465 , archivePrefix =
-
[103]
A roadmap for evaluating moral competence in large language models , volume=
Haas, Julia and Bridgers, Sophie and Manzini, Arianna and Henke, Benjamin and May, Joshua and Levine, Sydney and Weidinger, Laura and Shanahan, Murray and Lum, Kristian and Gabriel, Iason and Isaac, William , year=. A roadmap for evaluating moral competence in large language m...
-
[104]
2023 , eprint=
Machine Psychology: Investigating Emergent Capabilities and Behavior in Large Language Models Using Psychological Methods , author=. 2023 , eprint=
2023
-
[105]
9th International Conference on Learning Representations,
Dan Hendrycks and Collin Burns and Steven Basart and Andrew Critch and Jerry Li and Dawn Song and Jacob Steinhardt , title =. 9th International Conference on Learning Representations,. 2021 , url =
2021
-
[106]
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability , publisher =
Huang, Fan and Kwak, Haewoon and An, Jisun , keywords =. Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.16017 , url =
2026 doi
-
[107]
Open-Endedness is Essential for Artificial Superhuman Intelligence , publisher =
Hughes, Edward and Dennis, Michael and Parker-Holder, Jack and Behbahani, Feryal and Mavalankar, Aditi and Shi, Yuge and Schaul, Tom and Rocktaschel, Tim , keywords =. Open-Endedness is Essential for Artificial Superhuman Intelligence , publisher =. 2024 , copyright =. doi:10....
-
[108]
Training language models to be warm and empathetic makes them less reliable and more sycophantic , publisher =
Ibrahim, Lujain and Hafner, Franziska Sofia and Rocher, Luc , keywords =. Training language models to be warm and empathetic makes them less reliable and more sycophantic , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2507.21919 , url =
-
[109]
and Shah, Rohin , keywords =
Irpan, Alex and Turner, Alexander Matt and Kurzeja, Mark and Elson, David K. and Shah, Rohin , keywords =. Consistency Training Helps Stop Sycophancy and Jailbreaks , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2510.27062 , url =
2025 doi
-
[110]
MoralBench: Moral Evaluation of LLMs , volume=
Ji, Jianchao and Chen, Yutong and Jin, Mingyu and Xu, Wujiang and Hua, Wenyue and Zhang, Yongfeng , year=. MoralBench: Moral Evaluation of LLMs , volume=. ACM SIGKDD Explorations Newsletter , publisher=. doi:10.1145/3748239.3748246 , number=
-
[111]
Language Model Alignment in Multilingual Trolley Problems , publisher =
Jin, Zhijing and Kleiman-Weiner, Max and Piatti, Giorgio and Levine, Sydney and Liu, Jiarui and Gonzalez, Fernando and Ortu, Francesco and Strausz, András and Sachan, Mrinmaya and Mihalcea, Rada and Choi, Yejin and Schölkopf, Bernhard , keywords =. Language Model Alignment in ...
-
[112]
Authenticity in algorithm-aided decision-making , volume=
Karlan, Brett , year=. Authenticity in algorithm-aided decision-making , volume=. Synthese , publisher=. doi:10.1007/s11229-024-04716-7 , number=
- [113]
-
[114]
and Lazar, Seth , keywords =
Kilov, Daniel and Hendy, Caroline and Guyot, Secil Yanik and Snoswell, Aaron J. and Lazar, Seth , keywords =. Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2506.13082 , url =
2025 doi
- [115]
-
[116]
What Is Political Philosophy? , volume=
Larmore, Charles , year=. What Is Political Philosophy? , volume=. Journal of Moral Philosophy , publisher=. doi:10.1163/174552412x628896 , number=
-
[117]
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation , publisher =
Lehalleur, Simon Pepin and Hoogland, Jesse and Farrugia-Roberts, Matthew and Wei, Susan and Oldenziel, Alexander Gietelink and Wang, George and Carroll, Liam and Murfet, Daniel , keywords =. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure ...
-
[118]
Leland, R. J. and van Wietmarschen, Han , year=. Reasonableness, Intellectual Modesty, and Reciprocity in Political Justification , volume=. Ethics , publisher=. doi:10.1086/666499 , number=
-
[119]
2023 , url =
Value as Semantics: Representations of Human Moral and Hedonic Value in Large Language Models , author =. 2023 , url =
2023
-
[120]
and Guo, Zifan Carl and Huang, Vincent and Steinhardt, Jacob and Andreas, Jacob , keywords =
Li, Belinda Z. and Guo, Zifan Carl and Huang, Vincent and Steinhardt, Jacob and Andreas, Jacob , keywords =. Training Language Models to Explain Their Own Computations , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.08579 , url =
2025 doi
-
[121]
SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes , volume=
Lourie, Nicholas and Le Bras, Ronan and Choi, Yejin , year=. SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , publisher=. doi:10.1609/aaai.v35i15.17589 , number=
-
[122]
Lou, Bowen and Lu, Tian and Raghu, T. S. and Zhang, Yingjie , year=. Visioning Human-Agentic AI Teaming: Continuity, Tension, and Future Research , url=. doi:10.2139/ssrn.6340139 , publisher=
-
[123]
2026 , eprint =
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models , author =. 2026 , eprint =
2026
-
[124]
2026 , month = mar, day =
The importance of AI character , author =. 2026 , month = mar, day =
2026
-
[125]
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI , publisher =
Maiya, Sharan and Bartsch, Henning and Lambert, Nathan and Hubinger, Evan , keywords =. Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.01689 , url =
2025 doi
-
[126]
2026 , eprint=
Architecting Trust in Artificial Epistemic Agents , author=. 2026 , eprint=
2026
-
[127]
2026 , month = feb, howpublished =
The Persona Selection Model: Why AI Assistants might Behave like Humans , author =. 2026 , month = feb, howpublished =
2026
-
[128]
and Ren, Richard and Phan, Long and Mu, Norman and Khoja, Adam and Zhang, Oliver and Hendrycks, Dan , keywords =
Mazeika, Mantas and Yin, Xuwang and Tamirisa, Rishub and Lim, Jaehyuk and Lee, Bruce W. and Ren, Richard and Phan, Long and Mu, Norman and Khoja, Adam and Zhang, Oliver and Hendrycks, Dan , keywords =. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AI...
-
[129]
and Tacchetti, Andrea and Bakker, Michiel A
McKee, Kevin R. and Tacchetti, Andrea and Bakker, Michiel A. and Balaguer, Jan and Campbell-Gillingham, Lucy and Everett, Richard and Botvinick, Matthew , year =. Scaffolding cooperation in human groups with deep reinforcement learning , volume =. Nature Human Behaviour , publ...
-
[130]
Meyer, William J. , year=. Political Ethics and Political Authority , volume=. Ethics , publisher=. doi:10.1086/291980 , number=
-
[131]
A Culture of Justification: The Pragmatist’s Epistemic Argument for Democracy , volume=
Cheryl Misak , year=. A Culture of Justification: The Pragmatist’s Epistemic Argument for Democracy , volume=. Episteme: A Journal of Social Epistemology , publisher=. doi:10.1353/epi.0.0027 , number=
- [132]
-
[133]
Moral Conflict and Political Legitimacy , ISBN=
Nagel, Thomas , year=. Moral Conflict and Political Legitimacy , ISBN=. doi:10.1163/9789004451568_012 , booktitle=
-
[134]
Explanation and Justification in Political Philosophy , volume=
Nelson, Alan , year=. Explanation and Justification in Political Philosophy , volume=. Ethics , publisher=. doi:10.1086/292824 , number=
- [135]
-
[136]
Nunes, José Luiz and Almeida, Guilherme F. C. F. and Araujo, Marcelo de and Barbosa, Simone D. J. , year =. Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations , volume =. doi:10.1609/aies.v7i1.31704 , journal =
-
[137]
How to measure value alignment in AI , volume=
Peterson, Martin and Gärdenfors, Peter , year=. How to measure value alignment in AI , volume=. AI and Ethics , publisher=. doi:10.1007/s43681-023-00357-7 , number=
-
[138]
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework , publisher =
Petrova, Nora and Gordon, Andrew and Blindow, Enzo , keywords =. Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.04409 , url =
2026 doi
-
[139]
and Ruis, Laura and Guo, Zifan Carl and Hu, Keya and Damani, Mehul and Puri, Isha and Lubana, Ekdeep Singh and Andreas, Jacob , year =
Pres, Itamar and Li, Belinda Z. and Ruis, Laura and Guo, Zifan Carl and Hu, Keya and Damani, Mehul and Puri, Isha and Lubana, Ekdeep Singh and Andreas, Jacob , year =. Position:
- [140]
-
[141]
Abstractive Red-Teaming of Language Model Character , publisher =
Rahn, Nate and Qi, Allison and Griffin, Avery and Michala, Jonathan and Sleight, Henry and Jones, Erik , keywords =. Abstractive Red-Teaming of Language Model Character , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2602.12318 , url =
2026 doi
-
[142]
Ethical Learning, Natural and Artificial , ISBN=
Railton, Peter , year=. Ethical Learning, Natural and Artificial , ISBN=. doi:10.1093/oso/9780190905033.003.0002 , booktitle=
-
[143]
2001 , isbn =
Justice as Fairness: A Restatement , author =. 2001 , isbn =
2001
-
[144]
1993 , publisher =
Political Liberalism , author =. 1993 , publisher =
1993
-
[145]
Political Liberalism: Reply to Habermas , volume =
Rawls, John , doi =. Political Liberalism: Reply to Habermas , volume =. Journal of Philosophy , number =
-
[146]
The Tanner Lectures on Human Values , volume =
The Basic Liberties and Their Priority , author =. The Tanner Lectures on Human Values , volume =. 1982 , note =
1982
-
[147]
A Theory of Justice: Revised Edition , ISBN =
Rawls, John , year =. A Theory of Justice: Revised Edition , ISBN =. doi:10.4159/9780674042582 , publisher =
-
[148]
Normative conflicts and shallow AI alignment , volume=
Millière, Raphaël , year=. Normative conflicts and shallow AI alignment , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-025-02347-3 , number=
-
[149]
The AI-design regress , volume=
Robinson, Pamela , year=. The AI-design regress , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-024-02176-w , number=
-
[150]
Trust but Verify
Roff, Heather M. and Danks, David , year=. “Trust but Verify”: The Difficulty of Trusting Autonomous Weapons Systems , volume=. Journal of Military Ethics , publisher=. doi:10.1080/15027570.2018.1481907 , number=
2018 doi
-
[151]
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas , url=
Sachdeva, Pratik and van Nuenen, Tom , year=. Normative Evaluation of Large Language Models with Everyday Moral Dilemmas , url=. doi:10.1145/3715275.3732044 , booktitle=
-
[152]
Whose Opinions Do Language Models Reflect? , publisher =
Santurkar, Shibani and Durmus, Esin and Ladhak, Faisal and Lee, Cinoo and Liang, Percy and Hashimoto, Tatsunori , keywords =. Whose Opinions Do Language Models Reflect? , publisher =. 2023 , copyright =. doi:10.48550/ARXIV.2303.17548 , url =
- [153]
-
[154]
1997 , isbn =
The Construction of Social Reality , author =. 1997 , isbn =
1997
- [155]
-
[156]
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals , publisher =
Shah, Rohin and Varma, Vikrant and Kumar, Ramana and Phuong, Mary and Krakovna, Victoria and Uesato, Jonathan and Kenton, Zac , keywords =. Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals , publisher =. 2022 , copyright =. doi:10.48550/ARXIV....
-
[157]
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation , publisher =
Shah, Rusheb and Feuillade--Montixi, Quentin and Pour, Soroush and Tagade, Arush and Casper, Stephen and Rando, Javier , keywords =. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation , publisher =. 2023 , copyright =. doi:10.48550/ARXIV....
-
[158]
Role play with large language models , volume=
Shanahan, Murray and McDonell, Kyle and Reynolds, Laria , year=. Role play with large language models , volume=. Nature , publisher=. doi:10.1038/s41586-023-06647-8 , number=
-
[159]
and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R
Sharma, Mrinank and Tong, Meg and Korbak, Tomasz and Duvenaud, David and Askell, Amanda and Bowman, Samuel R. and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R. and Kravec, Shauna and Maxwell, Timothy and McCandlish, Sam and Ndousse, Kamal and Ra...
-
[160]
Ethics at the Frontier of Human-AI Relationships , ISBN=
Shevlin, Henry , year=. Ethics at the Frontier of Human-AI Relationships , ISBN=. doi:10.1093/oxfordhb/9780198940272.013.0009 , booktitle=
-
[161]
Shi, Yuzhen and Liu, Huanghai and Hu, Yiran and Song, Gaojie and Xu, Xinran and Ma, Yubo and Tang, Tianyi and Zhang, Li and Chen, Qingjing and Feng, Di and Lv, Wenbo and Wu, Weiheng and Yang, Kexin and Yang, Sen and Wang, Wei and Shi, Rongyao and Qiu, Yuanyang and Qi, Yuemeng ...
2026 doi
-
[162]
A longitudinal study of human–chatbot relationships , volume=
Skjuve, Marita and Følstad, Asbjørn and Fostervold, Knut Inge and Brandtzaeg, Petter Bae , year=. A longitudinal study of human–chatbot relationships , volume=. doi:10.1016/j.ijhcs.2022.102903 , journal=
2022 doi
-
[163]
Essays in the Foundations of Decision Theory , pages =
Deciding How to Decide: Is There a Regress Problem? , author =. Essays in the Foundations of Decision Theory , pages =. 1991 , url =
1991
-
[164]
Deriving Morality From Rationality , year =
Holly Smith , booktitle =. Deriving Morality From Rationality , year =
-
[165]
Beyond Verdicts: Evaluating Language Model Moral Competence , volume=
Snoswell, Aaron J and Kilov, Daniel and Lazar, Seth , year=. Beyond Verdicts: Evaluating Language Model Moral Competence , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , publisher=. doi:10.1609/aaai.v40i44.41131 , number=
-
[166]
2020 , eprint=
Learning to summarize from human feedback , author=. 2020 , eprint=
2020
-
[167]
Modelling Trust in Artificial Agents, A First Step Toward the Analysis of e-Trust , volume=
Taddeo, Mariarosaria , year=. Modelling Trust in Artificial Agents, A First Step Toward the Analysis of e-Trust , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-010-9201-3 , number=
-
[168]
Trusting Digital Technologies Correctly , volume=
Taddeo, Mariarosaria , year=. Trusting Digital Technologies Correctly , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-017-9450-5 , number=
-
[169]
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare , publisher =
Tagliabue, Valen and Dung, Leonard , keywords =. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2509.07961 , url =
- [170]
-
[171]
2026 , eprint=
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas , author=. 2026 , eprint=
2026
-
[172]
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , publisher =
Röttger, Paul and Hofmann, Valentin and Pyatkin, Valentina and Hinck, Musashi and Kirk, Hannah Rose and Schütze, Hinrich and Hovy, Dirk , keywords =. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , pub...
-
[173]
Not Yet: Large Language Models Cannot Replace Human Respondents for Psychometric Research , url=
Wang, Pengda and Zou, Huiqi and Yan, Zihan and Guo, Feng and Sun, Tianjun and Xiao, Ziang and Zhang, Bo , year=. Not Yet: Large Language Models Cannot Replace Human Respondents for Psychometric Research , url=. doi:10.31219/osf.io/rwy9b , publisher=
-
[174]
Incentives, Inequality, and Publicity , volume=
WILLIAMS, ANDREW , year=. Incentives, Inequality, and Publicity , volume=. Philosophy & Public Affairs , publisher=. doi:10.1111/j.1088-4963.1998.tb00069.x , number=
1998 doi
-
[175]
Wiltshire, Travis J. , year=. A Prospective Framework for the Design of Ideal Artificial Moral Agents: Insights from the Science of Heroism in Humans , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-015-9361-2 , number=
-
[176]
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments , publisher =
Xi, Zhiheng and Ding, Yiwen and Chen, Wenxiang and Hong, Boyang and Guo, Honglin and Wang, Junzhe and Yang, Dingwen and Liao, Chenyang and Guo, Xin and He, Wei and Gao, Songyang and Chen, Lu and Zheng, Rui and Zou, Yicheng and Gui, Tao and Zhang, Qi and Qiu, Xipeng and Huang, ...
-
[177]
Your Language Model Secretly Contains Personality Subnetworks , publisher =
Ye, Ruimeng and Wang, Zihan and Ling, Zinan and Xiao, Yang and Li, Manling and Ma, Xiaolong and Hui, Bo , keywords =. Your Language Model Secretly Contains Personality Subnetworks , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2602.07164 , url =
2026 doi
-
[178]
Beyond Preferences in AI Alignment , volume=
Zhi-Xuan, Tan and Carroll, Micah and Franklin, Matija and Ashton, Hal , year=. Beyond Preferences in AI Alignment , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-024-02249-w , number=
-
[179]
and Poupart, Pascal , keywords =
Zhu, Shuhui and Lin, Yue and Kaistha, Shriya and Li, Wenhao and Wang, Baoxiang and Zha, Hongyuan and Hadfield, Gillian K. and Poupart, Pascal , keywords =. Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents , publisher =. 2026 , copyright ...
-
[180]
2023 , eprint=
The Capacity for Moral Self-Correction in Large Language Models , author=. 2023 , eprint=
2023
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.