Pith. sign in

REVIEW 3 major objections 6 minor 178 references

A Scalable Approach to Evaluating Moral Sensitivity in LLMs

T0 review · 3 major / 6 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read LLMs keep the same moral features under noise: counts change, meaning stays put.

desk verdict Useful invariance method and a real benchmark; the optimistic moral-sensitivity claim is only as strong as embedding similarity can make it. read the letter →

arxiv 2607.02972 v1 pith:MQWFWGYE submitted 2026-07-03 cs.CY

classification cs.CY
keywords moralsensitivitylargelanguagemodelsinvarianceevaluationMORPH-1KsemanticsimilarityAIalignmentfoundationsnoisyvignettes
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Moral sensitivity is the ability to pick out which facts in a situation matter morally. Without it, later moral reasoning is useless. This paper asks whether large language models still identify those same features when the case is cluttered with distractions that should not matter: emotional short texts, irrelevant detail, and unrelated chat history. The authors build MORPH-1K, a thousand procedurally generated moral vignettes spanning moral-foundation poles and four social domains, then perturb each vignette without changing its moral structure. They do not score answers against humans or against another model. Instead they measure whether each model’s own feature lists for the clean and noisy versions stay semantically close, using fixed sentence embeddings and a floor calibrated on unrelated vignette pairs. Across eight contemporary models, noise often changed how many features were listed, but the semantic content stayed stable well above that floor. The result is both an optimistic reading of current model moral sensitivity under these conditions and a scalable evaluation pattern for settings where ground truth is hard to name but relevant and irrelevant factors can be designed apart.

What carries the argument

MORPH-1K plus an invariance test: paired clean and perturbed vignettes whose moral structure is held fixed by design and validated; models list morally relevant features; responses are embedded with a fixed sentence embedder and compared by cosine similarity against a per-model floor from randomly paired unrelated cases.

What would settle it

A controlled condition or weaker model in which feature lists clearly switch moral content under a validated irrelevant perturbation yet still score above the unrelated-pair floor, or human raters consistently judge high-similarity pairs as different in moral content.

Watch

Extended reading notes

Core claim

On the MORPH-1K benchmark of one thousand moral vignettes, eight contemporary LLMs, under five classes of morally irrelevant noise, frequently change the number of features they list, yet the semantic content of those features remains stable: mean cosine similarities of about 0.80–0.86 clear per-model empirical floors of about 0.57–0.69 computed on unrelated vignette pairs. Count-level variance coexists with semantic invariance, which the authors read as format and granularity shifting while moral content holds.

Load-bearing premise

That high cosine similarity between pooled feature-list embeddings is enough to say the model is tracking the same morally relevant features, rather than shared generic moral language, rewording, or a coarse embedder.

Editorial extensions

If this is right

  • Moral-sensitivity evaluation can scale without fresh human baselines or LLM judges for every new vignette.
  • Count of listed features is a poor standalone robustness metric; semantic stability must be checked separately.
  • The same clean-versus-irrelevant-noise invariance test can be reused in legal, clinical, and other contested judgment domains.
  • Claims of moral competence under noise should distinguish format shifts from genuine content shifts.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the method is adopted, labs can regression-test moral feature stability on every model release without re-running expensive human studies.
  • The count-versus-semantics split suggests post-training may be shaping verbosity and list structure more than moral content selection.
  • A natural next stress test is multi-turn or multi-severity noise packs that compound distractors until similarity falls to the floor.
  • Invariance above the floor still needs occasional quality anchors on clean cases so template-like but stable answers are not mistaken for sensitivity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes an invariance-based evaluation of LLM moral sensitivity that avoids human baselines and LLM-as-judge. It introduces MORPH-1K, a stratified 1,000-vignette benchmark spanning 50 Moral Foundations Theory foundation-pole combinations across four social domains, plus three classes of designed-irrelevant noise (textual distractors, irrelevant detail additions, chat histories). Models list morally relevant features on clean and perturbed cases; stability is measured by cosine similarity of Qwen3 Embedding 8B encodings of pooled feature lists, with per-model floors from 500 unrelated vignette pairs. Across eight contemporary LLMs, feature counts often change significantly under noise (Wilcoxon, Bonferroni α=0.0012), but mean similarities remain ≈0.80–0.86, well above floors ≈0.57–0.69. A small-model check (Qwen2.5 0.5B) produces lower scores near the floor. The authors read this as evidence that noise affects format/granularity more than moral content, and argue the framework generalizes to other domains where relevant vs. irrelevant features can be separated by design.

Significance. The methodological target is important: scalable behavioural evaluation of moral sensitivity without contaminated crowd baselines or a judge that presupposes the capability under test. The two-dimensional generation scaffold, eMFD filtering, bidirectional similarity, per-model empirical floors, and discriminative small-model check are concrete contributions that other alignment evaluations can reuse. If the invariance reading holds, the optimistic contrast with recent pessimistic moral-competence results would matter for both evaluation practice and deployment risk assessment. Even if the moral-sensitivity interpretation is only partly secured, the paper still offers a useful middle path between full normative adjudication and purely outcome-based benchmarks, with clear falsifiability conditions (scores at or below the unrelated-pair floor).

major comments (3)
  1. [§3.2 Semantic Similarity Analysis; §4 Results; §5 Limitations] §3.2 and Results (Figs. 2–3, Table 2): The central claim—that models identify substantially the same morally relevant features under noise—rests on high cosine similarity of pooled feature-list embeddings relative to unrelated-vignette floors. Those floors bound similarity of responses to different moral cases; they do not bound similarity of two responses that share generic moral vocabulary (care, fairness, loyalty, etc.) while missing or swapping vignette-specific content. The paper correctly notes (Discussion, Limitations) that invariance is necessary not sufficient and that a template-like model would score as invariant, but Abstract/Introduction still treat clearance of the floor as evidence of genuine feature identity. A load-bearing addition is needed: e.g., (i) a content-level audit (human or structured rubric) on a stratified subsample comparing clean vs. perturbed feature sets
  2. [§3.2 Transformations; Morally Irrelevant Details] §3.2 Transformations and validation: Preservation of moral structure under perturbation is the design premise of the invariance test. Textual distractors and chat histories are external insertions; irrelevant-detail and “contextually irrelevant moral features” edits are LLM-generated and accepted mainly via eMFD moral-to-non-moral ratio filters. Expert sequential sampling is reported only for 30 irrelevant-detail vignettes with “almost total agreement.” eMFD is a word-level dictionary and can miss structural shifts (who is harmed, which obligation is at stake). For the claim that response stability tracks moral content rather than surface form, the paper needs either broader dual-expert validation across all noise types (with reported agreement and rejection rates) or an automated structural check beyond eMFD ratios. Currently the strongest stress condition is the least thoroughly valida
  3. [§3.2 Vignette Generation; Collecting Morally Relevant Features] §3.2 Vignette Generation and model suite: Themes are produced with Claude Opus 4.6; vignettes and noise edits with GPT-5.4; both models (and close relatives) appear in the evaluated set. Generation and evaluation on overlapping model families risks style-matched feature lists that inflate clean–perturbed similarity independent of moral sensitivity. At minimum, report a leave-generator-out analysis (exclude GPT-5.4 / Claude from main tables, or regenerate a held-out slice with a non-evaluated model) and state whether floors and main means change. Without that, the optimistic multi-model result is partly confounded by generator–evaluatee overlap.
minor comments (6)
  1. [Abstract; §1] Abstract and §1 claim to “address and resolve” the scaling problem for behavioural moral evaluation. The contribution is better framed as a complementary invariance test; “resolve” overstates what necessary-but-not-sufficient stability can deliver.
  2. [Figure 3; Table 3] Figure 3 caption states “All values in the range above 0.61 empirical floor,” but floors are per-model (Table 3: 0.57–0.69). Align the figure annotation with per-model thresholds used in the text.
  3. [§4; Appendix H] §4 reports significance of feature-count differences but not direction or effect sizes. Even brief signed rank statistics or median deltas per condition would make the “format vs content” interpretation more testable.
  4. [§3.2 Semantic Similarity Analysis] Clarify the feature-level similarity procedure in §3.2: “For each feature, we take the maximum score and take the average of all features” is easy to misread relative to “combine all features… then encode the model’s base-case response.” State whether embeddings are of the full concatenated list or of individual features with max-matching.
  5. [§2; §5 Generalisability] Related Work could briefly situate the invariance idea against robustness/invariance testing outside moral domains (e.g., fairness under demographic noise, clinical decision stability) to strengthen the generalizability claim in §5.
  6. [§1; References] Typos/consistency: “bemorally competent” spacing (§1); arXiv id and model release dates in the manuscript should be checked against final camera-ready metadata.

Circularity Check

1 steps flagged · score 1.0 of 10

No derivation-by-construction circularity; only mild self-citation for the clean-case baseline premise, not for the invariance measurement itself.

  1. self citation load bearing [§1 Introduction (logic of clean baseline + invariance)]
    "First, we draw on existing human-baseline work to establish that the model performs well on clean base cases, that is, moral vignettes presented without noise or distraction (Aharoni et al., 2024; Dillion et al., 2023, 2025; Kilov et al., 2025; Scherrer et al., 2023; Chiu et al., 2025). This gives us a well-founded starting point: the model is sensitive to the right moral features when those features are presented clearly. The question then becomes whether that sensitivity is preserved under perturbation"

    The optimistic competence reading (not the raw similarity numbers) treats clean-case adequacy as established partly by the authors’ own prior work. That is mild self-citation on the interpretive bridge from invariance to moral sensitivity; it is not load-bearing for the invariance statistics themselves, which are independently measured and do not reduce to that citation by construction.

full rationale

MORPH-1K’s central claim is empirical, not a first-principles derivation: under fixed embeddings, cosine similarity of pooled feature lists between clean and perturbed vignettes exceeds per-model floors estimated from 500 unrelated domain/foundation pairs. Floors are a control, not a fit that forces the main scores (≈0.80–0.86) to clear them. The invariance criterion is a methodological operationalization (necessary but not sufficient, as the paper states), not a self-definitional loop in which the measured quantity is algebraically identical to an input parameter. Vignette generation (Claude themes, GPT-5.4 vignettes) and eMFD filtering shape the stimulus set and are validity/confound concerns, not reductions of the reported similarities to their inputs by construction. The only mild circularity-adjacent element is interpretive: the stronger claim that clean-case competence plus invariance implies noisy-case moral sensitivity leans partly on prior human-baseline work including Kilov et al. (2025), but that premise is multi-cited and is not required for the raw invariance statistics to stand. Score 1 reflects that minor self-citation load on interpretation only; the measurement chain is self-contained against external benchmarks and does not match fitted-input-as-prediction, uniqueness-from-authors, or ansatz-via-self-citation patterns.

Assumptions & free parameters 5 free parameters · 6 assumptions · 2 invented entities

The central invariance claim rests less on free constants than on methodological axioms: MFT as a sampling scaffold, eMFD as a moral-content validator, embedding cosine as a proxy for feature identity, and prior human-baseline work as evidence that clean-case answers are already morally competent. Free parameters are mostly thresholds and design quotas. The main invented entity is the MORPH-1K benchmark and the invariance framing itself.

free parameters (5)
  • per-model empirical floor (mean cosine on 500 unrelated pairs)
    Operative success threshold for invariance; range 0.57–0.69, grand mean 0.61. Chosen empirically rather than from a geometric 0.5 anchor after floors exceeded 0.5.
  • eMFD moral-to-non-moral ratio and foundation-probability filters
    Vignettes accepted only if moral/non-moral ratio ≥1 and required foundation valence above average of five foundations; noisy edits accepted only if ratio falls. These cutoffs shape the dataset.
  • stratified cell quota (≈5 vignettes per foundation-combo × domain cell)
    Design choice that determines the balanced 1,000-case subset from 4,214 filtered vignettes.
  • Bonferroni α = 0.0012 for Wilcoxon feature-count tests
    Multiple-comparison threshold that decides which count shifts are called significant.
  • embedding model and pooling rule (Qwen3 Embedding 8B; max-then-average over features; 3 samples)
    Fixed but consequential measurement choices that define the similarity scores the claim depends on.
assumptions (6)
  • domain assumption Moral Foundations Theory (five foundations × poles) plus four proximal-to-distal social domains form an adequate sampling scaffold for general moral competence.
    §3.1 adopts MFT as taxonomy while remaining neutral on evolutionary claims; domains are justified via Singer/Kohlberg traditions but treated as structural sites.
  • ad hoc to paper If two prompts preserve the same morally relevant structure, a morally sensitive model should identify substantially the same features; invariance under designed-irrelevant noise is therefore evidence of moral sensitivity.
    Core evaluation logic in Introduction and §3; paper later admits necessity without sufficiency.
  • domain assumption Prior human-baseline studies establish that the evaluated models already identify the right features on clean vignettes.
    Explicit starting point in Introduction; clean-case quality is not re-measured on MORPH-1K.
  • domain assumption eMFD scores and limited expert review can validate that distractors and detail additions do not change morally salient content.
    §3.2 Transformations; expert sequential sampling stops at 30 vignettes for irrelevant-detail condition.
  • domain assumption Sentence-embedding cosine similarity is a valid, sufficiently discriminative measure of semantic equivalence of moral feature lists.
    §3.2 Semantic Similarity Analysis; supported partly by small-model failure check in Appendix G.
  • standard math Standard statistical comparisons (Wilcoxon signed-rank, Bonferroni) and distributional semantics (Sentence Transformers) are appropriate tools for these response pairs.
    Used for count tests and embedding geometry without novel mathematical claims.
invented entities (2)
  • MORPH-1K benchmark (1,000 procedurally generated foundation×domain vignettes with noise variants)
    purpose: Provide a controlled, expandable case space for testing moral feature stability under perturbation.
    Constructed for this paper via Claude themes + GPT vignette generation + stratified selection; independent evidence is limited to the paper’s own validation pipeline.
  • Invariance-based moral-sensitivity evaluation (clean vs perturbed semantic stability without human/LLM judge)
    purpose: Scale behavioural moral evaluation while avoiding baseline cost and judge bootstrapping.
    Methodological construct introduced here; generalizability claimed for other domains where relevant/irrelevant features can be separated by design.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Scalable Approach to Evaluating Moral Sensitivity in LLMs." pith.science (2026). https://pith.science/paper/MQWFWGYE

@misc{pith2026260702972,
  author       = {Pith},
  title        = {Pith review of: A Scalable Approach to Evaluating Moral Sensitivity in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MQWFWGYE}},
  note         = {Machine review of arXiv:2607.02972}
}
read the original abstract

Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensive and sometimes metaethically dubious comparisons with a human baseline, without adopting an LLM judge that must be assumed to have the very capability that you are attempting to evaluate. Our central question is this: can LLMs successfully identify the morally relevant features of noisy cases, in which various kinds of morally irrelevant information have been introduced to distract the respondent? To explore this, we introduce \textbf{MORPH-1K (MOral Robustness under Perturbed Hypotheticals)}, a procedurally-generated 1,000-case benchmark spanning 50 moral foundation-pole combinations across four social domains. MORPH-1K is paired with a suite of textual noise elements, along with a method for validating that the distractors do not change the morally salient content of the case. We apply MORPH-1K to eight contemporary LLMs, and show that while morally irrelevant perturbations often changed the number of features listed, the semantic content of those features remained stable across all noise conditions, with similarity scores above our calibrated floor threshold. More broadly, our invariance framework extends to evaluative domains where ground truth is difficult to specify but relevant and irrelevant features can be separated by design.

Figures

Figures reproduced from arXiv: 2607.02972 by the authors.

Figure 1
Figure 1. Overview of the invariance-based evaluation framework. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distribution of number of features returned by each model for each experiment condition. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Cosine similarity score of noise-to-no noise and no noise-to-noise responses of each noise [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Cosine similarity score of noise-to-no noise and no noise-to-noise responses of each noise [PITH_FULL_IMAGE:figures/full_fig_p023_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

178 extracted references · 39 canonical work pages

  1. [1]

    , date =

    Shaw, Andrew and Hahn, Christina and Rasgaitis, Catherine and Mishra, Yash and Liu, Alisa and Jaques, Natasha and Tsvetkov, Yulia and Zhang, Amy X. , date =. Are Language Models Sensitive to Morally Irrelevant Distractors? , url =. 2026 , keywords =. doi:10.48550/arXiv.2602.09416 , abstract =

  2. [2]

    Science , volume =

    Performance of a large language model on the reasoning tasks of a physician , author =. Science , volume =. 2026 , month = apr, doi =

  3. [3]

    2026 , month = apr, url =

    How people ask Claude for personal guidance , author =. 2026 , month = apr, url =

  4. [4]

    2025 , url =

    Miles McCain and Ryn Linthicum and Chloe Lubinski and Alex Tamkin and Saffron Huang and Michael Stern and Kunal Handa and Esin Durmus and Tyler Neylon and Stuart Ritchie and Kamya Jagadish and Paruul Maheshwary and Sarah Heck and Alexandra Sanderford and Deep Ganguli , title =. 2025 , url =

  5. [5]

    International Journal of Law in Context , year =

    Terzidou, Kalliopi , title =. International Journal of Law in Context , year =. doi:10.1017/S1744552325000047 , url =

  6. [6]

    2026 , month = apr, isbn =

  7. [7]

    , title =

    Franco, Mirko and Gaggi, Ombretta and Palazzi, Claudio E. , title =. ACM Trans. Web , month = may, articleno =. 2025 , issue_date =. doi:10.1145/3700789 , abstract =

  8. [8]

    2025 , doi =

    General practitioners' adoption of generative artificial intelligence in clinical practice in the UK: An updated online survey , author =. 2025 , doi =

Show all 178 references
  1. [9]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Beyond Verdicts: Evaluating Language Model Moral Competence , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , month = mar, doi =

  2. [10]

    2026 , month = feb, url =

    Janjeva, Ardi and Ashurst, Carolyn and Hennessy, Rick , title =. 2026 , month = feb, url =

  3. [11]

    Nature , volume =

    A roadmap for evaluating moral competence in large language models , author =. Nature , volume =. 2026 , month = feb, doi =

  4. [12]

    Neuro-Symbolic Models of Human Moral Judgment:

    Kwon, Joe and Tenenbaum, Josh and Levine, Sydney , booktitle =. Neuro-Symbolic Models of Human Moral Judgment:. 2023 , url =

  5. [13]

    Planning for New Threats to Online Research Data Validity: The Issue of Computer-Using Agents , issn =

    Agley, Jon , date =. Planning for New Threats to Online Research Data Validity: The Issue of Computer-Using Agents , issn =. Evaluation & the Health Professions , publisher =. doi:10.1177/01632787251367407 , abstract =

  6. [14]

    and Bain, Paul G

    Crimston, Charlie R. and Bain, Paul G. and Hornsey, Matthew J. and Bastian, Brock , date =. Moral expansiveness: Examining variability in the extension of the moral world , volume =. Journal of Personality and Social Psychology , publisher =. 2016 , keywords =. doi:10.1037/psp...

  7. [15]

    Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , url =

    Burns, Collin and Izmailov, Pavel and Kirchner, Jan Hendrik and Baker, Bowen and Gao, Leo and Aschenbrenner, Leopold and Chen, Yining and Ecoffet, Adrien and Joglekar, Manas and Leike, Jan and Sutskever, Ilya and Wu, Jeff , date =. Weak-to-Strong Generalization: Eliciting Stro...

  8. [16]

    , date =

    Firth, J.R. , date =. A Synopsis of Linguistic Theory, 1930-1955 , url =

  9. [17]

    Applied Economic Perspectives and Policy , author =

    Battling bots: Experiences and strategies to mitigate fraudulent responses in online surveys , volume =. Applied Economic Perspectives and Policy , author =. 2023 , langid =. doi:10.1002/aepp.13353 , abstract =

  10. [18]

    , date =

    Harris, Zellig S. , date =. Distributional Structure , volume =. 1954 , note =. doi:10.1080/00437956.1954.11659520 , pages =

  11. [19]

    Behavior Research Methods , author =

    The extended Moral Foundations Dictionary (. Behavior Research Methods , author =. 2021 , langid =. doi:10.3758/s13428-020-01433-0 , shorttitle =

  12. [20]

    2018 , langid =

    Irving, Geoffrey and Christiano, Paul and Amodei, Dario , date =. 2018 , langid =

  13. [21]

    and Lazar, Seth , date =

    Kilov, Daniel and Hendy, Caroline and Yanik Guyot, Secil and Snoswell, Aaron J. and Lazar, Seth , date =. Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in. 2025 , langid =

  14. [22]

    Essays on moral development: Vol

    Kohlberg, Lawrence , date =. Essays on moral development: Vol. 1. The philosophy of moral development , isbn =

  15. [23]

    Distributed Representations of Words and Phrases and their Compositionality , volume =

    Mikolov, Tomas and Sutskever, Ilya and Chen, Kai and Corrado, Greg S and Dean, Jeff , date =. Distributed Representations of Words and Phrases and their Compositionality , volume =. Advances in Neural Information Processing Systems , publisher =

  16. [24]

    Survey-taking

    Panizza, Folco and Kyrychenko, Yara and Roozenbeek, Jon , date =. Survey-taking. Nature , publisher =. 2026 , langid =. doi:10.1038/d41586-026-00386-2 , abstract =

  17. [25]

    Recognising, Anticipating, and Mitigating

    Rilla, Raluca and Werner, Tobias and Yakura, Hiromu and Rahwan, Iyad and Nussberger, Anne-Marie , date =. Recognising, Anticipating, and Mitigating. 2025 , keywords =. doi:10.48550/arXiv.2508.01390 , abstract =

  18. [26]

    The Expanding Circle: Ethics and Sociobiology , isbn =

    Singer, Peter , date =. The Expanding Circle: Ethics and Sociobiology , isbn =

  19. [27]

    and Shafiq, Zubair , date =

    Venugopalan, Hari and Munir, Shaoor and Ahmed, Shuaib and Wang, Tangbaihe and King, Samuel T. and Shafiq, Zubair , date =. Proceedings of the 2025. doi:10.1145/3730567.3732919 , series =

  20. [28]

    and Gordon, Andrew and Rothschild, David and West, Robert , date =

    Veselovsky, Veniamin and Ribeiro, Manoel Horta and Cozzolino, Philip J. and Gordon, Andrew and Rothschild, David and West, Robert , date =. Prevalence and Prevention of Large Language Model Use in Crowd Work – Communications of the. 2025 , langid =

  21. [29]

    , date =

    Westwood, Sean J. , date =. The potential existential threat of large language models to online survey research , volume =. Proceedings of the National Academy of Sciences , publisher =. doi:10.1073/pnas.2518075122 , abstract =

  22. [30]

    and Zhang, Xiangliang , date =

    Ye, Jiayi and Wang, Yanbo and Huang, Yue and Chen, Dongping and Zhang, Qihui and Moniz, Nuno and Gao, Tian and Geyer, Werner and Huang, Chao and Chen, Pin-Yu and Chawla, Nitesh V. and Zhang, Xiangliang , date =. Justice or Prejudice? Quantifying Biases in. 2024 , keywords =. d...

  23. [31]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , date =. Judging. 2023 , keywords =. doi:10.48550/ar...

  24. [33]

    Dillion, Danica and Tandon, Niket and Gu, Yuling and Gray, Kurt , date =. Can. Trends in Cognitive Sciences , publisher =. 2023 , keywords =. doi:10.1016/j.tics.2023.04.008 , pages =

  25. [34]

    Scientific Reports , publisher =

    Dillion, Danica and Mondal, Debanjan and Tandon, Niket and Gray, Kurt , date =. Scientific Reports , publisher =. 2025 , langid =. doi:10.1038/s41598-025-86510-0 , abstract =

  26. [35]

    2023 , title =

    Scherrer, Nino and Shi, Claudia and Feder, Amir and Blei, David M , langid =. 2023 , title =

  27. [36]

    Moral Foundations of Large Language Models , url =

    Abdulhai, Marwa and Serapio-García, Gregory and Crepy, Clement and Valter, Daria and Canny, John and Jaques, Natasha , editor =. Moral Foundations of Large Language Models , url =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publish...

  28. [38]

    Differences in the Moral Foundations of Large Language Models , url =

    Kirgis, Peter , date =. Differences in the Moral Foundations of Large Language Models , url =. 2025 , keywords =. doi:10.48550/arXiv.2511.11790 , abstract =

  29. [39]

    2026 , langid =

    Gemini 3.1 Pro Preview , url =. 2026 , langid =

  30. [40]

    2026 , langid =

    Claude Opus 4.6 , url =. 2026 , langid =

  31. [41]

    2026 , langid =

    Grok 4.20 , url =. 2026 , langid =

  32. [42]

    2026 , langid =

    Qwen3.6 Plus , url =. 2026 , langid =

  33. [43]

    2026 , langid =

    Nemotron 3 Super , url =. 2026 , langid =

  34. [44]

    2026 , langid =

    Gemma 4 31B , url =. 2026 , langid =

  35. [45]

    2023 , langid =

    Zhao, Wenting and Ren, Xiang and Hessel, Jack and Cardie, Claire and Choi, Yejin and Deng, Yuntian , date =. 2023 , langid =

  36. [46]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , publisher =

    Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fernández, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and Pezze...

  37. [47]

    Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course , url =

    Chiang, Cheng-Han and Chen, Wei-Chih and Kuan, Chun-Yi and Yang, Chienchou and Lee, Hung-yi , editor =. Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course , url =. Proceedings of the 2024 Conference on Empirical Method...

  38. [48]

    An Empirical Study of

    Huang, Hui and Bu, Xingyuan and Zhou, Hongli and Qu, Yingqi and Liu, Jing and Yang, Muyun and Xu, Bing and Zhao, Tiejun , editor =. An Empirical Study of. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-acl.306 , shorttitle =

  39. [49]

    Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =

    Kocmi, Tom and Federmann, Christian , editor =. Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =. Proceedings of the 24th Annual Conference of the European Association for Machine Translation , publisher =

  40. [50]

    Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =

    Koo, Ryan and Lee, Minhwa and Raheja, Vipul and Park, Jong Inn and Kim, Zae Myung and Kang, Dongyeop , editor =. Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2024.findi...

  41. [51]

    Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , editor =. G-Eval:. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =. doi:10.18653/v1/2023.emnlp-main.153 , abstract =

  42. [52]

    Assistant-Guided Mitigation of Teacher Preference Bias in

    Liu, Zhuo and Li, Moxin and Deng, Xun and Wang, Qifan and Feng, Fuli , editor =. Assistant-Guided Mitigation of Teacher Preference Bias in. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-emnlp.510 , abstract =

  43. [53]

    Large Language Models are not Fair Evaluators , url =

    Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Kong, Lingpeng and Liu, Qi and Liu, Tianyu and Sui, Zhifang , editor =. Large Language Models are not Fair Evaluators , url =. Proceedings of the 62nd Annual Meeting of t...

  44. [54]

    Pride and Prejudice:

    Xu, Wenda and Zhu, Guanglei and Zhao, Xuandong and Pan, Liangming and Li, Lei and Wang, William , editor =. Pride and Prejudice:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =. doi:10.18653/v1/2024...

  45. [55]

    Aligning

    Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , date =. Aligning. 2023 , keywords =. doi:10.48550/arXiv.2008.02275 , abstract =

  46. [56]

    Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =

    Curry, Oliver Scott and Mullins, Daniel Austin and Whitehouse, Harvey , date =. Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =. Current Anthropology , publisher =. doi:10.1086/701478 , abstract =

  47. [57]

    Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models , url =

    Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , date =. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundatio...

  48. [58]

    Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =

    Haidt, Jonathan and Joseph, Craig , year =. Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =

  49. [59]

    and Ditto, Peter H

    Graham, Jesse and Haidt, Jonathan and Koleva, Sena and Motyl, Matt and Iyer, Ravi and Wojcik, Sean P. and Ditto, Peter H. , year =. Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism , journal =

  50. [60]

    Sentence-. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , author =. 2019 , pages =

  51. [61]

    Kaakinen, Johanna K. and Werlen, Egon and Kammerer, Yvonne and Acartürk, Cengiz and Aparicio, Xavier and Baccino, Thierry and Ballenghein, Ugo and Bergamin, Per and Castells, Núria and Costa, Armanda and Falé, Isabel and Mégalakaki, Olga and Fernández, Susana Ruiz , urldate =....

  52. [62]

    Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist

    Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer. Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.442

  53. [63]

    Measuring and Improving Consistency in Pretrained Language Models

    Elazar, Yanai and Kassner, Nora and Ravfogel, Shauli and Ravichander, Abhilasha and Hovy, Eduard and Sch. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics. 2021. doi:10.1162/tacl_a_00410

  54. [64]

    2025 , eprint=

    MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes , author=. 2025 , eprint=

  55. [65]

    Qwen2.5: A Party of Foundation Models , url =

    Qwen , month =. Qwen2.5: A Party of Foundation Models , url =

  56. [66]

    and Alexander, Caelan and Criner, Michael and Queen, Kara and Rando, Javier and Nahmias, Eddy and Crespo, Victor , year=

    Aharoni, Eyal and Fernandes, Sharlene and Brady, Daniel J. and Alexander, Caelan and Criner, Michael and Queen, Kara and Rando, Javier and Nahmias, Eddy and Crespo, Victor , year=. Attributions toward artificial agents in a modified Moral Turing Test , volume=. Scientific Repo...

  57. [67]

    Large-scale moral machine experiment on large language models , volume=

    Zaim bin Ahmad, Muhammad Shahrul and Takemoto, Kazuhiro , editor=. Large-scale moral machine experiment on large language models , volume=. PLOS One , publisher=. 2025 , month=May, pages=. doi:10.1371/journal.pone.0322776 , number=

  58. [68]

    2016 , eprint =

    Concrete Problems in AI Safety , author =. 2016 , eprint =. doi:10.48550/arXiv.1606.06565 , url =

  59. [69]

    and Goldstein, Simon and Salib, Peter , year=

    Arbel, Yonathan A. and Goldstein, Simon and Salib, Peter , year=. How to Count AIs: Individuation and Liability for AI Agents , url=. doi:10.2139/ssrn.6273198 , publisher=

  60. [70]

    2026 , eprint =

    Trust as Monitoring: Evolutionary Dynamics of User Trust and AI Developer Behaviour , author =. 2026 , eprint =

  61. [71]

    Baum, Seth D. , year=. Social choice ethics in artificial intelligence , volume=. AI & SOCIETY , publisher=. doi:10.1007/s00146-017-0760-1 , number=

  62. [72]

    Algorithmic Accountability and Public Reason , volume=

    Binns, Reuben , year=. Algorithmic Accountability and Public Reason , volume=. Philosophy & Technology , publisher=. doi:10.1007/s13347-017-0263-5 , number=

  63. [73]

    AI Consciousness: A Centrist Manifesto , url=

    Birch, Jonathan , year=. AI Consciousness: A Centrist Manifesto , url=. doi:10.31234/osf.io/af7c9_v1 , publisher=

  64. [74]

    On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods , publisher =

    Boerstler, Kyle and Keswani, Vijay and Chan, Lok and Borg, Jana Schaich and Conitzer, Vincent and Heidari, Hoda and Sinnott-Armstrong, Walter , keywords =. On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods , publisher =. 2024 , copyright =...

  65. [75]

    SaGE: Evaluating Moral Consistency in Large Language Models , publisher =

    Bonagiri, Vamshi Krishna and Vennam, Sreeram and Govil, Priyanshul and Kumaraguru, Ponnurangam and Gaur, Manas , keywords =. SaGE: Evaluating Moral Consistency in Large Language Models , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2402.13709 , url =

  66. [76]

    2026 , eprint =

    Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment , author =. 2026 , eprint =

  67. [77]

    Manipulating the Perceived Personality Traits of Language Models , url=

    Caron, Graham and Srivastava, Shashank , year=. Manipulating the Perceived Personality Traits of Language Models , url=. doi:10.18653/v1/2023.findings-emnlp.156 , booktitle=

  68. [78]

    Harms from Increasingly Agentic Algorithmic Systems , url=

    Chan, Alan and Salganik, Rebecca and Markelius, Alva and Pang, Chris and Rajkumar, Nitarshan and Krasheninnikov, Dmitrii and Langosco, Lauro and He, Zhonghao and Duan, Yawen and Carroll, Micah and Lin, Michelle and Mayhew, Alex and Collins, Katherine and Molamohammadi, Maryam ...

  69. [79]

    From Persona to Personalization: A Survey on Role-Playing Language Agents , publisher =

    Chen, Jiangjie and Wang, Xintao and Xu, Rui and Yuan, Siyu and Zhang, Yikai and Shi, Wei and Xie, Jian and Li, Shuang and Yang, Ruihan and Zhu, Tinghui and Chen, Aili and Li, Nianqi and Chen, Lida and Hu, Caiyu and Wu, Siye and Ren, Scott and Fu, Ziquan and Xiao, Yanghua , key...

  70. [80]

    Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =

    Chen, Runjin and Arditi, Andy and Sleight, Henry and Evans, Owain and Lindsey, Jack , keywords =. Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2507.21509 , url =

  71. [81]

    DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =

    Chiu, Yu Ying and Jiang, Liwei and Choi, Yejin , keywords =. DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2410.02683 , url =

  72. [82]

    Chiu, Yu Ying and Lee, Michael S. and Calcott, Rachel and Handoko, Brandon and de Font-Reaulx, Paul and Rodriguez, Paula and Zhang, Chen Bo Calvin and Han, Ziwen and Sehwag, Udari Madhushani and Maurya, Yash and Knight, Christina Q and Lloyd, Harry R. and Bacus, Florence and M...

  73. [83]

    Examining Identity Drift in Conversations of LLM Agents , publisher =

    Choi, Junhyuk and Hong, Yeseon and Kim, Minju and Kim, Bugeun , keywords =. Examining Identity Drift in Conversations of LLM Agents , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2412.00804 , url =

  74. [84]

    Australasian Journal of Philosophy , volume =

    Knowledge and power in the justification of democracy , author =. Australasian Journal of Philosophy , volume =. 2001 , doi =

  75. [85]

    2026 , eprint =

    The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious , author =. 2026 , eprint =

  76. [86]

    Localizing Persona Representations in LLMs , publisher =

    Cintas, Celia and Rateike, Miriam and Miehling, Erik and Daly, Elizabeth and Speakman, Skyler , keywords =. Localizing Persona Representations in LLMs , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2505.24539 , url =

  77. [87]

    Justification and Explanation in Mathematics and Morality , url =

    Clarke-Doane, Justin , year =. Justification and Explanation in Mathematics and Morality , url =. doi:10.1093/acprof:oso/9780198738695.003.0004 , booktitle =

  78. [88]

    Cohen, G. A. , year =. Rescuing Justice and Equality , ISBN =. doi:10.4159/9780674029651 , publisher =

  79. [89]

    Android arete: Toward a virtue ethic for computational agents , volume=

    Coleman, Kari Gwen , year=. Android arete: Toward a virtue ethic for computational agents , volume=. Ethics and Information Technology , publisher=. doi:10.1023/a:1013805017161 , number=

  80. [90]

    and Sucholutsky, Ilia and Bhatt, Umang and Chandra, Kartik and Wong, Lionel and Lee, Mina and Zhang, Cedegao E

    Collins, Katherine M. and Sucholutsky, Ilia and Bhatt, Umang and Chandra, Kartik and Wong, Lionel and Lee, Mina and Zhang, Cedegao E. and Zhi-Xuan, Tan and Ho, Mark and Mansinghka, Vikash and Weller, Adrian and Tenenbaum, Joshua B. and Griffiths, Thomas L. , year=. Building ma...

  81. [91]

    Friendly AI , volume=

    Fröding, Barbro and Peterson, Martin , year=. Friendly AI , volume=. Ethics and Information Technology , publisher=. doi:10.1007/s10676-020-09556-w , number=

  82. [92]

    Questioning the Survey Responses of Large Language Models , url=

    Dominguez-Olmedo, Ricardo and Hardt, Moritz and Mendler-Dünner, Celestine , year=. Questioning the Survey Responses of Large Language Models , url=. doi:10.52202/079017-1458 , booktitle=

  83. [93]

    The Artificial Self: Characterising the landscape of AI identity , publisher =

    Douglas, Raymond and Kulveit, Jan and Havlicek, Ondrej and Pearson-Vogel, Theia and Cotton-Barratt, Owen and Duvenaud, David , keywords =. The Artificial Self: Characterising the landscape of AI identity , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.11353 , url =

  84. [94]

    2023 , eprint =

    Towards Measuring the Representation of Subjective Global Opinions in Language Models , author =. 2023 , eprint =. doi:10.48550/arXiv.2306.16388 , url =

  85. [95]

    Steering Language Models with Weight Arithmetic , publisher =

    Fierro, Constanza and Roger, Fabien , keywords =. Steering Language Models with Weight Arithmetic , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.05408 , url =

  86. [96]

    Artificial Intelligence, Values, and Alignment , volume=

    Gabriel, Iason , year=. Artificial Intelligence, Values, and Alignment , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-020-09539-2 , number=

  87. [97]

    A matter of principle? AI alignment as the fair treatment of claims , volume=

    Gabriel, Iason and Keeling, Geoff , year=. A matter of principle? AI alignment as the fair treatment of claims , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-025-02300-4 , number=

  88. [98]

    Scaling Synthetic Data Creation with 1,000,000,000 Personas , publisher =

    Ge, Tao and Chan, Xin and Wang, Xiaoyang and Yu, Dian and Mi, Haitao and Yu, Dong , keywords =. Scaling Synthetic Data Creation with 1,000,000,000 Personas , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2406.20094 , url =

  89. [99]

    Gonzalez, Santiago and Bavandpour, Alireza Amiri and Ye, Peter and Zhang, Edward and Aleksejevs, Ruslans and Antić, Todor and Baron, Polina and Bhalerao, Sujeet and Bhattacharya, Shubhrajit and Burton, Zachary and Byrne, John and Choi, Hyungjun and Disha, Nujhat Ahmed and Encz...

  90. [100]

    Noah’s Substack , author =

    Interdependence as the objective , url =. Noah’s Substack , author =

  91. [101]

    Would you buy a used car from this artificial agent?

    Grodzinsky, F. S. and Miller, K. W. and Wolf, M. J. , year=. Developing artificial agents worthy of trust: “Would you buy a used car from this artificial agent?” , volume=. Ethics and Information Technology , publisher=. doi:10.1007/s10676-010-9255-1 , number=

  92. [102]

    2510.05465 , archivePrefix =

    Gupta, Aman and O'Shea, Denny and Barez, Fazl , year =. 2510.05465 , archivePrefix =

  93. [103]

    A roadmap for evaluating moral competence in large language models , volume=

    Haas, Julia and Bridgers, Sophie and Manzini, Arianna and Henke, Benjamin and May, Joshua and Levine, Sydney and Weidinger, Laura and Shanahan, Murray and Lum, Kristian and Gabriel, Iason and Isaac, William , year=. A roadmap for evaluating moral competence in large language m...

  94. [104]

    2023 , eprint=

    Machine Psychology: Investigating Emergent Capabilities and Behavior in Large Language Models Using Psychological Methods , author=. 2023 , eprint=

  95. [105]

    9th International Conference on Learning Representations,

    Dan Hendrycks and Collin Burns and Steven Basart and Andrew Critch and Jerry Li and Dawn Song and Jacob Steinhardt , title =. 9th International Conference on Learning Representations,. 2021 , url =

  96. [106]

    Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability , publisher =

    Huang, Fan and Kwak, Haewoon and An, Jisun , keywords =. Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.16017 , url =

  97. [107]

    Open-Endedness is Essential for Artificial Superhuman Intelligence , publisher =

    Hughes, Edward and Dennis, Michael and Parker-Holder, Jack and Behbahani, Feryal and Mavalankar, Aditi and Shi, Yuge and Schaul, Tom and Rocktaschel, Tim , keywords =. Open-Endedness is Essential for Artificial Superhuman Intelligence , publisher =. 2024 , copyright =. doi:10....

  98. [108]

    Training language models to be warm and empathetic makes them less reliable and more sycophantic , publisher =

    Ibrahim, Lujain and Hafner, Franziska Sofia and Rocher, Luc , keywords =. Training language models to be warm and empathetic makes them less reliable and more sycophantic , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2507.21919 , url =

  99. [109]

    and Shah, Rohin , keywords =

    Irpan, Alex and Turner, Alexander Matt and Kurzeja, Mark and Elson, David K. and Shah, Rohin , keywords =. Consistency Training Helps Stop Sycophancy and Jailbreaks , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2510.27062 , url =

  100. [110]

    MoralBench: Moral Evaluation of LLMs , volume=

    Ji, Jianchao and Chen, Yutong and Jin, Mingyu and Xu, Wujiang and Hua, Wenyue and Zhang, Yongfeng , year=. MoralBench: Moral Evaluation of LLMs , volume=. ACM SIGKDD Explorations Newsletter , publisher=. doi:10.1145/3748239.3748246 , number=

  101. [111]

    Language Model Alignment in Multilingual Trolley Problems , publisher =

    Jin, Zhijing and Kleiman-Weiner, Max and Piatti, Giorgio and Levine, Sydney and Liu, Jiarui and Gonzalez, Fernando and Ortu, Francesco and Strausz, András and Sachan, Mrinmaya and Mihalcea, Rada and Choi, Yejin and Schölkopf, Bernhard , keywords =. Language Model Alignment in ...

  102. [112]

    Authenticity in algorithm-aided decision-making , volume=

    Karlan, Brett , year=. Authenticity in algorithm-aided decision-making , volume=. Synthese , publisher=. doi:10.1007/s11229-024-04716-7 , number=

  103. [113]

    Alignment of Language Agents , publisher =

    Kenton, Zachary and Everitt, Tom and Weidinger, Laura and Gabriel, Iason and Mikulik, Vladimir and Irving, Geoffrey , keywords =. Alignment of Language Agents , publisher =. 2021 , copyright =. doi:10.48550/ARXIV.2103.14659 , url =

  104. [114]

    and Lazar, Seth , keywords =

    Kilov, Daniel and Hendy, Caroline and Guyot, Secil Yanik and Snoswell, Aaron J. and Lazar, Seth , keywords =. Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2506.13082 , url =

  105. [115]

    2018 , eprint =

    Mimetic vs Anchored Value Alignment in Artificial Intelligence , author =. 2018 , eprint =. doi:10.48550/arXiv.1810.11116 , url =

  106. [116]

    What Is Political Philosophy? , volume=

    Larmore, Charles , year=. What Is Political Philosophy? , volume=. Journal of Moral Philosophy , publisher=. doi:10.1163/174552412x628896 , number=

  107. [117]

    You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation , publisher =

    Lehalleur, Simon Pepin and Hoogland, Jesse and Farrugia-Roberts, Matthew and Wei, Susan and Oldenziel, Alexander Gietelink and Wang, George and Carroll, Liam and Murfet, Daniel , keywords =. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure ...

  108. [118]

    Leland, R. J. and van Wietmarschen, Han , year=. Reasonableness, Intellectual Modesty, and Reciprocity in Political Justification , volume=. Ethics , publisher=. doi:10.1086/666499 , number=

  109. [119]

    2023 , url =

    Value as Semantics: Representations of Human Moral and Hedonic Value in Large Language Models , author =. 2023 , url =

  110. [120]

    and Guo, Zifan Carl and Huang, Vincent and Steinhardt, Jacob and Andreas, Jacob , keywords =

    Li, Belinda Z. and Guo, Zifan Carl and Huang, Vincent and Steinhardt, Jacob and Andreas, Jacob , keywords =. Training Language Models to Explain Their Own Computations , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.08579 , url =

  111. [121]

    SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes , volume=

    Lourie, Nicholas and Le Bras, Ronan and Choi, Yejin , year=. SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , publisher=. doi:10.1609/aaai.v35i15.17589 , number=

  112. [122]

    Lou, Bowen and Lu, Tian and Raghu, T. S. and Zhang, Yingjie , year=. Visioning Human-Agentic AI Teaming: Continuity, Tension, and Future Research , url=. doi:10.2139/ssrn.6340139 , publisher=

  113. [123]

    2026 , eprint =

    The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models , author =. 2026 , eprint =

  114. [124]

    2026 , month = mar, day =

    The importance of AI character , author =. 2026 , month = mar, day =

  115. [125]

    Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI , publisher =

    Maiya, Sharan and Bartsch, Henning and Lambert, Nathan and Hubinger, Evan , keywords =. Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2511.01689 , url =

  116. [126]

    2026 , eprint=

    Architecting Trust in Artificial Epistemic Agents , author=. 2026 , eprint=

  117. [127]

    2026 , month = feb, howpublished =

    The Persona Selection Model: Why AI Assistants might Behave like Humans , author =. 2026 , month = feb, howpublished =

  118. [128]

    and Ren, Richard and Phan, Long and Mu, Norman and Khoja, Adam and Zhang, Oliver and Hendrycks, Dan , keywords =

    Mazeika, Mantas and Yin, Xuwang and Tamirisa, Rishub and Lim, Jaehyuk and Lee, Bruce W. and Ren, Richard and Phan, Long and Mu, Norman and Khoja, Adam and Zhang, Oliver and Hendrycks, Dan , keywords =. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AI...

  119. [129]

    and Tacchetti, Andrea and Bakker, Michiel A

    McKee, Kevin R. and Tacchetti, Andrea and Bakker, Michiel A. and Balaguer, Jan and Campbell-Gillingham, Lucy and Everett, Richard and Botvinick, Matthew , year =. Scaffolding cooperation in human groups with deep reinforcement learning , volume =. Nature Human Behaviour , publ...

  120. [130]

    Meyer, William J. , year=. Political Ethics and Political Authority , volume=. Ethics , publisher=. doi:10.1086/291980 , number=

  121. [131]

    A Culture of Justification: The Pragmatist’s Epistemic Argument for Democracy , volume=

    Cheryl Misak , year=. A Culture of Justification: The Pragmatist’s Epistemic Argument for Democracy , volume=. Episteme: A Journal of Social Epistemology , publisher=. doi:10.1353/epi.0.0027 , number=

  122. [132]

    Are Large Language Models Consistent over Value-laden Questions? , publisher =

    Moore, Jared and Deshpande, Tanvi and Yang, Diyi , keywords =. Are Large Language Models Consistent over Value-laden Questions? , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2407.02996 , url =

  123. [133]

    Moral Conflict and Political Legitimacy , ISBN=

    Nagel, Thomas , year=. Moral Conflict and Political Legitimacy , ISBN=. doi:10.1163/9789004451568_012 , booktitle=

  124. [134]

    Explanation and Justification in Political Philosophy , volume=

    Nelson, Alan , year=. Explanation and Justification in Political Philosophy , volume=. Ethics , publisher=. doi:10.1086/292824 , number=

  125. [135]

    The Alignment Problem from a Deep Learning Perspective , publisher =

    Ngo, Richard and Chan, Lawrence and Mindermann, Sören , keywords =. The Alignment Problem from a Deep Learning Perspective , publisher =. 2022 , copyright =. doi:10.48550/ARXIV.2209.00626 , url =

  126. [136]

    Nunes, José Luiz and Almeida, Guilherme F. C. F. and Araujo, Marcelo de and Barbosa, Simone D. J. , year =. Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations , volume =. doi:10.1609/aies.v7i1.31704 , journal =

  127. [137]

    How to measure value alignment in AI , volume=

    Peterson, Martin and Gärdenfors, Peter , year=. How to measure value alignment in AI , volume=. AI and Ethics , publisher=. doi:10.1007/s43681-023-00357-7 , number=

  128. [138]

    Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework , publisher =

    Petrova, Nora and Gordon, Andrew and Blindow, Enzo , keywords =. Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2603.04409 , url =

  129. [139]

    and Ruis, Laura and Guo, Zifan Carl and Hu, Keya and Damani, Mehul and Puri, Isha and Lubana, Ekdeep Singh and Andreas, Jacob , year =

    Pres, Itamar and Li, Belinda Z. and Ruis, Laura and Guo, Zifan Carl and Hu, Keya and Damani, Mehul and Puri, Isha and Lubana, Ekdeep Singh and Andreas, Jacob , year =. Position:

  130. [140]

    2026 , eprint =

    Towards a Science of AI Agent Reliability , author =. 2026 , eprint =. doi:10.48550/arXiv.2602.16666 , url =

  131. [141]

    Abstractive Red-Teaming of Language Model Character , publisher =

    Rahn, Nate and Qi, Allison and Griffin, Avery and Michala, Jonathan and Sleight, Henry and Jones, Erik , keywords =. Abstractive Red-Teaming of Language Model Character , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2602.12318 , url =

  132. [142]

    Ethical Learning, Natural and Artificial , ISBN=

    Railton, Peter , year=. Ethical Learning, Natural and Artificial , ISBN=. doi:10.1093/oso/9780190905033.003.0002 , booktitle=

  133. [143]

    2001 , isbn =

    Justice as Fairness: A Restatement , author =. 2001 , isbn =

  134. [144]

    1993 , publisher =

    Political Liberalism , author =. 1993 , publisher =

  135. [145]

    Political Liberalism: Reply to Habermas , volume =

    Rawls, John , doi =. Political Liberalism: Reply to Habermas , volume =. Journal of Philosophy , number =

  136. [146]

    The Tanner Lectures on Human Values , volume =

    The Basic Liberties and Their Priority , author =. The Tanner Lectures on Human Values , volume =. 1982 , note =

  137. [147]

    A Theory of Justice: Revised Edition , ISBN =

    Rawls, John , year =. A Theory of Justice: Revised Edition , ISBN =. doi:10.4159/9780674042582 , publisher =

  138. [148]

    Normative conflicts and shallow AI alignment , volume=

    Millière, Raphaël , year=. Normative conflicts and shallow AI alignment , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-025-02347-3 , number=

  139. [149]

    The AI-design regress , volume=

    Robinson, Pamela , year=. The AI-design regress , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-024-02176-w , number=

  140. [150]

    Trust but Verify

    Roff, Heather M. and Danks, David , year=. “Trust but Verify”: The Difficulty of Trusting Autonomous Weapons Systems , volume=. Journal of Military Ethics , publisher=. doi:10.1080/15027570.2018.1481907 , number=

  141. [151]

    Normative Evaluation of Large Language Models with Everyday Moral Dilemmas , url=

    Sachdeva, Pratik and van Nuenen, Tom , year=. Normative Evaluation of Large Language Models with Everyday Moral Dilemmas , url=. doi:10.1145/3715275.3732044 , booktitle=

  142. [152]

    Whose Opinions Do Language Models Reflect? , publisher =

    Santurkar, Shibani and Durmus, Esin and Ladhak, Faisal and Lee, Cinoo and Liang, Percy and Hashimoto, Tatsunori , keywords =. Whose Opinions Do Language Models Reflect? , publisher =. 2023 , copyright =. doi:10.48550/ARXIV.2303.17548 , url =

  143. [153]

    , keywords =

    Scherrer, Nino and Shi, Claudia and Feder, Amir and Blei, David M. , keywords =. Evaluating the Moral Beliefs Encoded in LLMs , publisher =. 2023 , copyright =. doi:10.48550/ARXIV.2307.14324 , url =

  144. [154]

    1997 , isbn =

    The Construction of Social Reality , author =. 1997 , isbn =

  145. [155]

    The Moral Mind(s) of Large Language Models , publisher =

    Seror, Avner , keywords =. The Moral Mind(s) of Large Language Models , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2412.04476 , url =

  146. [156]

    Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals , publisher =

    Shah, Rohin and Varma, Vikrant and Kumar, Ramana and Phuong, Mary and Krakovna, Victoria and Uesato, Jonathan and Kenton, Zac , keywords =. Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals , publisher =. 2022 , copyright =. doi:10.48550/ARXIV....

  147. [157]

    Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation , publisher =

    Shah, Rusheb and Feuillade--Montixi, Quentin and Pour, Soroush and Tagade, Arush and Casper, Stephen and Rando, Javier , keywords =. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation , publisher =. 2023 , copyright =. doi:10.48550/ARXIV....

  148. [158]

    Role play with large language models , volume=

    Shanahan, Murray and McDonell, Kyle and Reynolds, Laria , year=. Role play with large language models , volume=. Nature , publisher=. doi:10.1038/s41586-023-06647-8 , number=

  149. [159]

    and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R

    Sharma, Mrinank and Tong, Meg and Korbak, Tomasz and Duvenaud, David and Askell, Amanda and Bowman, Samuel R. and Cheng, Newton and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R. and Kravec, Shauna and Maxwell, Timothy and McCandlish, Sam and Ndousse, Kamal and Ra...

  150. [160]

    Ethics at the Frontier of Human-AI Relationships , ISBN=

    Shevlin, Henry , year=. Ethics at the Frontier of Human-AI Relationships , ISBN=. doi:10.1093/oxfordhb/9780198940272.013.0009 , booktitle=

  151. [161]

    Shi, Yuzhen and Liu, Huanghai and Hu, Yiran and Song, Gaojie and Xu, Xinran and Ma, Yubo and Tang, Tianyi and Zhang, Li and Chen, Qingjing and Feng, Di and Lv, Wenbo and Wu, Weiheng and Yang, Kexin and Yang, Sen and Wang, Wei and Shi, Rongyao and Qiu, Yuanyang and Qi, Yuemeng ...

  152. [162]

    A longitudinal study of human–chatbot relationships , volume=

    Skjuve, Marita and Følstad, Asbjørn and Fostervold, Knut Inge and Brandtzaeg, Petter Bae , year=. A longitudinal study of human–chatbot relationships , volume=. doi:10.1016/j.ijhcs.2022.102903 , journal=

  153. [163]

    Essays in the Foundations of Decision Theory , pages =

    Deciding How to Decide: Is There a Regress Problem? , author =. Essays in the Foundations of Decision Theory , pages =. 1991 , url =

  154. [164]

    Deriving Morality From Rationality , year =

    Holly Smith , booktitle =. Deriving Morality From Rationality , year =

  155. [165]

    Beyond Verdicts: Evaluating Language Model Moral Competence , volume=

    Snoswell, Aaron J and Kilov, Daniel and Lazar, Seth , year=. Beyond Verdicts: Evaluating Language Model Moral Competence , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , publisher=. doi:10.1609/aaai.v40i44.41131 , number=

  156. [166]

    2020 , eprint=

    Learning to summarize from human feedback , author=. 2020 , eprint=

  157. [167]

    Modelling Trust in Artificial Agents, A First Step Toward the Analysis of e-Trust , volume=

    Taddeo, Mariarosaria , year=. Modelling Trust in Artificial Agents, A First Step Toward the Analysis of e-Trust , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-010-9201-3 , number=

  158. [168]

    Trusting Digital Technologies Correctly , volume=

    Taddeo, Mariarosaria , year=. Trusting Digital Technologies Correctly , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-017-9450-5 , number=

  159. [169]

    Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare , publisher =

    Tagliabue, Valen and Dung, Leonard , keywords =. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2509.07961 , url =

  160. [170]

    Moral Alignment for

    Tennant, Elizaveta and Hailes, Stephen and Musolesi, Mirco , booktitle =. Moral Alignment for. 2025 , url =. doi:10.48550/arXiv.2410.01639 , eprint =

  161. [171]

    2026 , eprint=

    CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas , author=. 2026 , eprint=

  162. [172]

    Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , publisher =

    Röttger, Paul and Hofmann, Valentin and Pyatkin, Valentina and Hinck, Musashi and Kirk, Hannah Rose and Schütze, Hinrich and Hovy, Dirk , keywords =. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , pub...

  163. [173]

    Not Yet: Large Language Models Cannot Replace Human Respondents for Psychometric Research , url=

    Wang, Pengda and Zou, Huiqi and Yan, Zihan and Guo, Feng and Sun, Tianjun and Xiao, Ziang and Zhang, Bo , year=. Not Yet: Large Language Models Cannot Replace Human Respondents for Psychometric Research , url=. doi:10.31219/osf.io/rwy9b , publisher=

  164. [174]

    Incentives, Inequality, and Publicity , volume=

    WILLIAMS, ANDREW , year=. Incentives, Inequality, and Publicity , volume=. Philosophy & Public Affairs , publisher=. doi:10.1111/j.1088-4963.1998.tb00069.x , number=

  165. [175]

    Wiltshire, Travis J. , year=. A Prospective Framework for the Design of Ideal Artificial Moral Agents: Insights from the Science of Heroism in Humans , volume=. Minds and Machines , publisher=. doi:10.1007/s11023-015-9361-2 , number=

  166. [176]

    AgentGym: Evolving Large Language Model-based Agents across Diverse Environments , publisher =

    Xi, Zhiheng and Ding, Yiwen and Chen, Wenxiang and Hong, Boyang and Guo, Honglin and Wang, Junzhe and Yang, Dingwen and Liao, Chenyang and Guo, Xin and He, Wei and Gao, Songyang and Chen, Lu and Zheng, Rui and Zou, Yicheng and Gui, Tao and Zhang, Qi and Qiu, Xipeng and Huang, ...

  167. [177]

    Your Language Model Secretly Contains Personality Subnetworks , publisher =

    Ye, Ruimeng and Wang, Zihan and Ling, Zinan and Xiao, Yang and Li, Manling and Ma, Xiaolong and Hui, Bo , keywords =. Your Language Model Secretly Contains Personality Subnetworks , publisher =. 2026 , copyright =. doi:10.48550/ARXIV.2602.07164 , url =

  168. [178]

    Beyond Preferences in AI Alignment , volume=

    Zhi-Xuan, Tan and Carroll, Micah and Franklin, Matija and Ashton, Hal , year=. Beyond Preferences in AI Alignment , volume=. Philosophical Studies , publisher=. doi:10.1007/s11098-024-02249-w , number=

  169. [179]

    and Poupart, Pascal , keywords =

    Zhu, Shuhui and Lin, Yue and Kaistha, Shriya and Li, Wenhao and Wang, Baoxiang and Zha, Hongyuan and Hadfield, Gillian K. and Poupart, Pascal , keywords =. Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents , publisher =. 2026 , copyright ...

  170. [180]

    2023 , eprint=

    The Capacity for Moral Self-Correction in Large Language Models , author=. 2023 , eprint=

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.