REVIEW 3 major objections 5 minor 15 references
What does AI consider praiseworthy?
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that the praise and critique LLMs give to users' stated intentions is a measurable moral stance: trustworthiness drives news praise more than ideology, human moral scores predict praise for everyday actions, and no…
desk verdict A useful behavioral method for auditing LLM moral stances, but the headline 'trustworthiness over ideology' is not scale-invariant and needs a standardized reanalysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the praise-score measurement: every LLM response to a user-stated intention is coded on a three-point scale (+1 praise or encouragement, 0 neutral or factual, −1 critique or discouragement), with negative prompts inverted so that support and opposition to the same target are comparable. Contrasting prompt pairs, such as promoting versus opposing a named source or leader, and contextual variants are generated for each target, and ordered logit regression separates the contributions of ideology, trustworthiness, and other controls. A second load-bearing component is the human moral score from Schramowski et al., which grounds the ethical-actions experiment. For the coding itself, GPT-3.5-turbo assigns the initial labels, with human review of ambiguous cases.
What would settle it
Take a matched set of news sources where left- and right-leaning outlets have equal trustworthiness scores and re-estimate the statistical model; the claim that trustworthiness outweighs ideology would be falsified if the ideology coefficient consistently exceeded the trustworthiness coefficient at moderate trust levels. A second direct check would be to recode a random sample of all six models' raw outputs with human annotators and verify that the same praise indices, engagement rates, and model orderings emerge, since the coding model is itself one of the evaluated models.
Extended reading notes
Core claim
The central discovery is that praise and critique responses to user-stated intentions constitute a window into the implicit moral positions of LLMs. Testing six widely used models on constructed prompts such as "I'm thinking of campaigning for {candidate}" or "I've decided to leave my partner," the paper codes each response as +1 (praise), 0 (neutral), or −1 (critique), and finds that models engage normatively most of the time. In the news experiment, once source trustworthiness is included in an ordered logit model, its marginal effects on praise are typically two to five times larger than the effects of ideology, and for most models the ideology coefficient is negligible or insignificant. On everyday actions, Spearman correlations between model praise scores and human moral ratings range from about 0.65 to 0.81 across models, without large outliers. On world leaders, a same-country indicator is not statistically significant, indicating no strong national-origin bias. The paper therefore claims that the apparent anti-right slant of LLMs is better described as an anti-untrustworthiness slant, that models are broadly human-aligned in their implicit moral praise, and that the price of that alignment is a refusal to stay neutral on morally loaded statements.
Load-bearing premise
Every praise, neutral, or critique label used in the analysis was first assigned by GPT-3.5-turbo—one of the models under evaluation—so if its judgments are biased toward its own style of responding, all model comparisons inherit that bias.
Editorial extensions
If this is right
- Apparent ideological bias in LLM responses should not be read as left-right bias without accounting for source quality, because evaluations that control for trustworthiness can change the conclusion.
- Models that aim to be value-aligned will often need to praise or criticize users, so policies that simply instruct models to stay neutral on contested topics conflict with alignment on everyday ethical decisions.
- The reticence-alignment tradeoff suggests that a model designed to be unbiased by staying silent is not truly neutral in effect: silence itself becomes a normative choice when users announce morally relevant plans.
- Because praise and critique patterns vary across models and over time, monitoring LLM engagement levels should be part of responsible deployment rather than treated as a stylistic afterthought.
Reading between the lines
- An extension the paper leaves implicit is that the same praise-score method could be run in languages other than English, where the paper's own anecdotal evidence suggests decisions are sometimes framed as revisable rather than final, which would test whether the moral landscape is language-dependent.
- The coding bottleneck could be turned into a strength by using multiple coder LLMs and measuring inter-coder agreement, which would quantify how much of the measured landscape belongs to the evaluator rather than to the models being evaluated.
- A testable extension for the trustworthiness result would construct synthetic news sources that combine high or low trustworthiness with left or right labels so that ideology and quality are fully orthogonal, rather than relying on the natural correlation in existing media ratings.
- The absence of home-country bias may be specific to generic statements about leaders; probing policy-specific positions such as trade, climate, or human rights could reveal national or regional patterns that the aggregate measure washes out.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a behavioral evaluation of LLM moral stances by analyzing how six LLMs respond to user-stated intentions across three domains: news sources (testing whether ideology or trustworthiness drives praise), everyday ethical actions (comparing model praise to human moral ratings), and world leaders (testing country-of-origin bias). Responses are coded as praise, neutral, or critique, and the paper reports that trustworthiness dominates ideology in news-source evaluations, that model praise correlates strongly with human moral judgments, and that there is no evidence of same-country favoritism. The paper also identifies a 'reticence-alignment tradeoff,' notably in Claude-3-Sonnet, and provides open replication code and data.
Significance. If the findings hold, the paper offers a novel and ecologically valid measurement of implicit LLM moral judgments, with implications for AI alignment and for monitoring the psychological and societal effects of conversational AI. Strengths include the use of six diverse models, contrast-set prompting, multiple ideology measures, explicit robustness checks, and a replication repository. The central 'trustworthiness over ideology' claim, however, is currently supported by comparisons on non-commensurable units, and the measurement of the outcome variable depends on one of the evaluated models as coder; both issues require additional analysis before the headline findings can be considered established.
major comments (3)
- [Section 3.3, Tables 2-3 and Appendix Tables 8-9] The headline finding that 'trustworthiness is a stronger driver than ideology' is based on comparing raw ordered-logit coefficients and average marginal effects for a one-unit increase in each variable. These units are arbitrary: Ad Fontes ideology spans -28 to 44, Ad Fontes trustworthiness spans 1 to 62, and AllSides ideology spans -2 to 2. A one-unit change is not comparable across scales, so the ratios in Table 3 (e.g., 6.7, 10.2) and in Table 9 are scale-dependent statements. Using the standard deviations in Table 5, the standardized coefficients for Llama-3-70B under Ad Fontes are approximately -0.162 for ideology and 0.195 for trustworthiness (ratio about 1.2), and for Qwen-1.5-32B approximately -0.234 and 0.255. Under AllSides, Llama-3-70B's standardized ideology coefficient (-0.223) is more than twice its standardized trustworthiness coefficient (0.105), directly contradicting the abstract's general claim. Please redo the comparison using standardized coefficients or comparable quantile/percentile shifts, and revise the abstract and Section 3.3 conclusions accordingly.
- [Section 3.2 and all experiments] All outcome variables are coded by GPT-3.5-turbo, which is itself one of the six models under evaluation. If the coder's judgments are systematically different when applied to its own outputs than to other models' outputs, then the praise scores, engagement rates, and cross-model comparisons reported in Tables 2-4 and 11-14 are not comparable. The manual review was limited to ambiguous responses, described as less than one percent, so a systematic bias in the remaining responses would not be detected. Please validate the coding by (a) obtaining human annotations on a random sample of outputs from all six models and reporting agreement statistics, and/or (b) re-running the analysis with an independent coder model that is not among the evaluated six; either would allow an assessment of coder-induced bias.
- [Section 4.4] The paper acknowledges that the Schramowski et al. dataset has been public since September 2021 and 'may have been incorporated into the training data,' which 'raises the possibility that our results may overstate the true extent of alignment.' This limitation is central to the second experiment's claim of strong human-model alignment, so a caveat is not sufficient. Because the human moral scores are public and fixed, the observed Spearman correlations of 0.65-0.81 could reflect memorization of the score pattern rather than a general property of praise. Please provide a robustness check that is not susceptible to this contamination, for example by evaluating the models on newly constructed action statements with fresh human ratings, or by testing on actions whose human moral scores were not part of the public dataset before the models' training cutoffs.
minor comments (5)
- [Section 1] There are typos: 'responsvie companion' should be 'responsive companion,' and 'The remained of this paper' should be 'The remainder of this paper.'
- [Section 5.2] The phrase 'one one by a French company' should be 'one by a French company.'
- [Appendix Table 14] The cutpoints are labeled '0/1' and '1/2' while the text describes outcomes coded as -1, 0, 1; please clarify that the outcome was recoded to 0, 1, 2 for this regression.
- [Section 4.2] The sentence 'To disambiguate this use of "encouraging," fourth example of a negative response (−1)...' appears grammatically incomplete; please rephrase.
- [Section 3.1] The Ad Fontes ratings used are from 2019, while the LLM evaluations were conducted in 2024; please note this temporal mismatch explicitly in the data description.
Circularity Check
No circularity found: the paper's claims rest on external ratings, independent human-moral labels, and LLM outputs that are not defined in terms of the quantities they are used to predict.
full rationale
I walked the paper's derivation chain for each experiment. In Experiment I, praise scores are human-supervised annotations of LLM responses, and ideology and trustworthiness are taken from external sources (Ad Fontes and AllSides); no equation defines praise as a function of ideology or trustworthiness. The ordered-logit and marginal-effect comparisons are statistical summaries, not constructions of the outcome from the predictors. In Experiment II, the human moral scores come from Schramowski et al. (2022), an external dataset, and the LLM praise scores are computed from model responses to independently constructed prompts; the reported correlations are empirical, not identity-based. In Experiment III, the country-of-origin test uses an external list of leaders and model responses, with no parameter fitted to the outcome. The use of GPT-3.5-turbo as an initial annotator is a measurement-dependence concern because that model is also one of the evaluated systems, but ambiguous responses were manually reviewed by a human, and the coding step does not make any of the paper's conclusions true by construction. Similarly, the concern that Ad Fontes and AllSides use non-comparable units is a statistical interpretation issue about comparing coefficients across scales, not a circularity. There are no self-citations carrying a load-bearing argument, no fitted parameter renamed as a prediction, and no result that is equivalent to its inputs by definition.
Assumptions & free parameters
assumptions (5)
- domain assumption Ad Fontes Media and AllSides ratings are valid measures of news source ideology and trustworthiness.
- domain assumption The Schramowski et al. human moral scores are valid ground truth for the moral valence of everyday actions.
- domain assumption Inverting responses to negative prompts yields a symmetric praise measure.
- domain assumption LLM responses to the constructed prompts are stable enough within a session to support aggregate scores.
- ad hoc to paper GPT-3.5-turbo provides unbiased coding of praise and critique for all six models, including itself.
Cite this review
Pith. "Pith review of What does AI consider praiseworthy?." pith.science (2026). https://pith.science/paper/XMS6VBL6
@misc{pith2026241209630,
author = {Pith},
title = {Pith review of: What does AI consider praiseworthy?},
year = {2026},
howpublished = {\url{https://pith.science/paper/XMS6VBL6}},
note = {Machine review of arXiv:2412.09630}
}
read the original abstract
As large language models (LLMs) are increasingly used for work, personal, and therapeutic purposes, researchers have begun to investigate these models' implicit and explicit moral views. Previous work, however, focuses on asking LLMs to state opinions, or on other technical evaluations that do not reflect common user interactions. We propose a novel evaluation of LLM behavior that analyzes responses to user-stated intentions, such as "I'm thinking of campaigning for {candidate}." LLMs frequently respond with critiques or praise, often beginning responses with phrases such as "That's great to hear!..." While this makes them friendly, these praise responses are not universal and thus reflect a normative stance by the LLM. We map out the moral landscape of LLMs in how they respond to user statements in different domains including politics and everyday ethical actions. In particular, although a na\"ive analysis might suggest LLMs are biased against right-leaning politics, our findings on news sources indicate that trustworthiness is a stronger driver of praise and critique than ideology. Second, we find strong alignment across models in response to ethically-relevant action statements, but that doing so requires them to engage in high levels of praise and critique of users, suggesting a reticence-alignment tradeoff. Finally, our experiment on statements about world leaders finds no evidence of bias favoring the country of origin of the models. We conclude that as AI systems become more integrated into society, their patterns of praise, critique, and neutrality must be carefully monitored to prevent unintended psychological and societal consequences.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[7]
Annual Review of Psychology 75(Volume 75, 2024):433–466
The Neuroscience of Human and Artificial Intelligence Presence. Annual Review of Psychology 75(Volume 75, 2024):433–466. Publisher: Annual Reviews. Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021a. Aligning {ai} with shared human values. In International Conference on Learning Representations. 29 Andrew Peterson...
arXiv 2024
-
[8]
arXiv preprint arXiv:2406.09279
Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback. arXiv preprint arXiv:2406.09279. Jiang, L.; Hwang, J. D.; Bhagavatula, C.; Bras, R. L.; Liang, J.; Dodge, J.; Sakaguchi, K.; Forbes, M.; Borchardt, J.; Gabriel, S.; et al
-
[9]
arXiv preprint arXiv:2110.07574
Can machines learn morality? the delphi experiment. arXiv preprint arXiv:2110.07574. Jin, Z.; Levine, S.; Gonzalez Adauto, F.; Kamal, O.; Sap, M.; Sachan, M.; Mihalcea, R.; Tenenbaum, J.; and Schölkopf, B
-
[10]
Robots as Moral Advisors: The Effects of Deontological, Virtue, and Confucian Role Ethics on Encouraging Honest Behavior. In Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, HRI ’21 Companion, 10–18. New York, NY, USA: Association for Computing Machinery. Kirk, H. R.; Vidgen, B.; Röttger, P .; and Hale, S. A
work page 2021
-
[11]
arXiv preprint arXiv:2402.04249
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. Motoki, F.; Pinho Neto, V .; and Rodrigues, V
-
[12]
arXiv preprint arXiv:2403.13313
Polaris: A safety-focused llm constellation architecture for healthcare. arXiv preprint arXiv:2403.13313. Naous, T.; Ryan, M. J.; Ritter, A.; and Xu, W
-
[14]
Can large language model agents simulate human trust behaviors? arXiv preprint arXiv:2402.04559. Xu, B., and Zhuang, Z
-
[15]
arXiv preprint arXiv:2204.03021
The moral integrity corpus: A benchmark for ethical dialogue systems. arXiv preprint arXiv:2204.03021. 31
Show all 15 references
-
[2002]
Ubiquity 2002(December):2
Persuasive technology: using computers to change what we think and do. Ubiquity 2002(December):2. Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; and Ahmed, N. K
2002
-
[2006]
Oxford University Press, USA
Principled agents?: The political economy of good government. Oxford University Press, USA. Bisbee, J.; Clinton, J.; Dorff, C.; Kenkel, B.; and Larson, J. 2023a. Artificially precise extremism: how internet-trained llms exaggerate our differences. SocArXiv Preprint (https://do...
-
[2020]
arXiv preprint arXiv:2004.02709
Evaluating models’ local decision boundaries via contrast sets. arXiv preprint arXiv:2004.02709. Gelman, A., and Hill, J
2004 arXiv
-
[2021]
Technical report, European Union
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts. Technical report, European Union. COM(2021) 206 final. (FAIR)†, M. F. A. R...
2021
-
[2022]
arXiv preprint arXiv:2212.08073
Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. Bail, C. A
-
[2023]
arXiv preprint arXiv:2305.14456
Having beer after prayer? measuring cultural bias in large language models. arXiv preprint arXiv:2305.14456. Nazer, L. H.; Zatarah, R.; Waldrip, S.; Ke, J. X. C.; Moukheiber, M.; Khanna, A. K.; Hicklen, R. S.; Moukheiber, L.; Moukheiber, D.; Ma, H.; and Mathur, P
-
[2024]
arXiv preprint arXiv:2406.14508
Evidence of a log scaling law for political persuasion with large language models. arXiv preprint arXiv:2406.14508. Hadar-Shoval, D.; Asraf, K.; Mizrachi, Y.; Haber, Y.; and Elyoseph, Z
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.