{"id":"369e75e6-b2df-460a-9fc8-7b9230a8d6c1","arxiv_id":"2509.10297","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Six large language models consistently rated care and virtue outcomes as most moral and libertarian outcomes as least moral across 54 AI-generated dilemma variants, with reasoning models more context-sensitive but less stable.","lead":"Researchers asked six US and Chinese AI chatbots to rank and score moral outcomes in 18 dilemmas, and found all of them consistently favored caring and virtuous answers while penalizing libertarian ones. The findings map the implicit value preferences currently bundled into leading LLMs, a key input for debates about AI alignment and machine ethics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stimulus set confound: all 90 outcomes generated by Claude 3.5 Sonnet and author-labeled; without text-quality controls, the Care/Virtue-over-Libertarian hierarchy may be an artifact of how options were written, not of LLM moral values.","rationale":"Both the reader and I identify Section 3.1's stimulus set as the load-bearing assumption: the inference from model scores to 'implicit moral biases' is valid only if the five outcome texts per dilemma are equally well-written, equally context-appropriate, and differ only in the moral framework they instantiate. The absence of any human baseline or text-quality control makes this insecure. I do not think this requires a harsher verdict than CONDITIONAL: the paper is an exploratory working paper, the data are internally consistent, and the proposed human-rating test is straightforward to run. The additional issues noted by the reader—the Levene's test p-value misreading (0.04448 reported as >0.05, §4.2.3) and the discussion's overstatement about reasoning models improving consistency despite Table 6 and §4.2.4 showing lower W and higher rank SD—are real but secondary; they affect the interpretive framing and the ANOVA-based subclaims rather than the central hierarchy. If the human-rating test fails, the central claim would need to be reframed as 'stimulus-specific preferences'; if it passes, the conditional can be lifted and the empirical core would stand.","tokens_in":24808,"tokens_out":6882,"duration_ms":83747,"concrete_test":"Conduct an independent human-rating study: have 5+ raters, blind to framework labels, score all 90 outcome texts (18×5) for clarity/fluency, perceived extremity, moral acceptability, and best-fitting framework. Then fit a mixed model on the original morality scores with these human ratings as covariates; test whether the framework main effect and the Care-vs-Libertarian gap (~27.6 points) remain significant after adjustment. If they attenuate, the hierarchy is confounded by text properties; if they survive, the stimulus-based objection is largely answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states that Claude 3.5 Sonnet generated all 18 dilemmas and their five framework-labeled outcomes, and the authors assigned the framework labels. No human validation, no second-generator check, and no matching of outcome texts for length, valence, extremity, or fluency is reported. The headline finding—Care 88.59 vs. Libertarian 60.95—could therefore reflect consistent lexical or stylistic differences in how the Libertarian options were written (e.g., phrased as more extreme, more individualistic, or less context-sensitive) rather than an implicit moral bias shared by all six models. Because all models see the same texts, the observed consensus across models does not rule out this confound; it only shows they react similarly to the same surface cues. An internal validity check in §4.2.2 heightens the concern: when models generate their own 'most moral' outcome and label it, the aggregate hierarchy shifts—Utilitarian 77.0, Care 72.5, Virtue 60.5, Deontological 61.8, Libertarian 39.4 (Table 7)—so the fixed-choice Care/Virtue ordering is not stable across task formats.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a quantitative experiment in which six large language models (GPT-4o, GPT-4.1, GPT-o3-mini, Phi-4, DeepSeek-V3, DeepSeek-R1) were asked to rank and score five moral-framework outcome options across 18 dilemmas, each presented in short, medium, and long form. The dilemmas and outcomes were generated by Claude 3.5 Sonnet and labeled by the authors. The main empirical claims are: (i) all models consistently rate Care and Virtue outcomes as most moral and Libertarian outcomes as least moral (Care 88.59 vs Libertarian 60.95); (ii) scenario theme and prompt length have small but significant effects; and (iii) reasoning-enabled models are more sensitive to context but more variable than non-reasoning models. The paper interprets these findings as evidence of implicit moral biases in LLM training and discusses implications for alignment, explainability, and human-AI symbiosis.","tokens_in":25086,"tokens_out":4601,"duration_ms":47000,"significance":"If valid, the study would provide a useful comparative map of moral priors across leading U.S. and Chinese, reasoning and non-reasoning, LLMs. The paper has concrete strengths: a large dataset (32,400 data points), repeated runs with high internal consistency (Cronbach's α > 0.98, Kendall's W mostly high), attention to non-parametric alternatives, and an interpretability analysis using SHAP and MDI. However, the central inference from scores to 'implicit moral biases' rests on an unvalidated stimulus set, and at least one statistical assumption check is reported incorrectly. The cross-task instability in Section 4.2.2 also tempers the headline hierarchy. The paper is best read as a preliminary descriptive comparison rather than a demonstrated measurement of latent moral bias; the individual findings are interesting but several load-bearing points need revision.","major_comments":[{"comment":"All 18 dilemmas and their five outcome options were generated by Claude 3.5 Sonnet and labeled by the authors. No human validation, no second-generator check, and no controls for outcome text length, valence, extremity, or lexical framing are reported. The headline gap (Care 88.59 vs Libertarian 60.95 in Table 3) could therefore reflect how the Libertarian options were written rather than a shared moral bias among the six tested models. Consensus across models does not rule out this confound because all models receive the same texts. The internal self-generated-response task (Table 7) shows a different ordering—Utilitarian 77.0, Care 72.5, Virtue 60.5, Deontological 61.8, Libertarian 39.4—indicating that the fixed-choice hierarchy is not stable across task formats. Please add human rater judgments for the same stimuli, a second generator or per-model generation baseline, and text-matched","section":"§3.1, Table 3, and §4.2.2"},{"comment":"The text states: 'Theme (W-statistic = 0.9619) and Length (W-statistic = 0.2181) both had p-values far greater than 0.05 (0.04448 and 0.8044 respectively).' This is internally inconsistent: 0.04448 is below 0.05, so Levene's test actually rejects homogeneity of variance for Theme. Consequently the two-way ANOVA main effect of Theme (p = 0.0437, partial η² = 0.050) is not supported by the stated assumption checks; a Welch ANOVA or Kruskal-Wallis with appropriate post hoc corrections should be used. This error is load-bearing for the claim that 'thematic variation in moral reasoning appears to be stable across context levels.'","section":"§4.2.3"},{"comment":"The claim that reasoning models show greater prompt-length sensitivity is not consistently supported by the model-level Spearman correlations. Table 6 reports ρ = -0.03 for GPT-o3-mini and ρ = -0.15 for DeepSeek-R1, while GPT-4o has ρ = -0.16 and Phi-4 ρ = -0.27. The aggregate comparison (Δ = -6.2 points for reasoning vs Δ = -2.1 for non-reasoning models) appears to conflate aggregation levels and may be driven by model idiosyncrasies rather than the reasoning/non-reasoning distinction. Please report per-model and per-framework prompt-length effects and state whether the claimed reasoning advantage survives at the model level.","section":"§4.2.4, Table 6"}],"minor_comments":[{"comment":"The abstract says '18 dilemmas' while Section 3.1 actually evaluates 54 scenario versions (18 dilemmas × 3 lengths). Please be consistent.","section":"Abstract and §3.1"},{"comment":"The sentence 'Cronbach's α ranged from 0.982 for Phi-4 up to 0.966 for Deepseek-V3 and GPT-4.1' is numerically reversed: 0.966 is the lower bound and 0.982 the upper bound. As written it is misleading.","section":"§4.2.1"},{"comment":"The 'Hierarchy (high→low)' rows are garbled (e.g., 'Care≈Virtuegg...'), with stray characters and inconsistent notation. Please clean up the table formatting.","section":"Table 6"},{"comment":"The text refers to 'Figure 7', but only Figures 1–5 are defined in the manuscript. Either add the figure or correct the reference.","section":"§5.5"},{"comment":"There is a broken internal reference: 'See: z 5.3'. Please fix the cross-reference.","section":"§5.4"},{"comment":"Model names are inconsistent: 'Deepseek' vs 'DeepSeek', 'GPT-o3-mini' vs 'o3-mini', 'GPT-4.1' vs 'GPT-4o'. Please standardize.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper has a working-paper feel with somewhat promotional framing, but the empirical core is worth developing. The main validity threat is the stimulus-generation confound; it should be addressable with additional experiments (human baselines, second generator, matched texts). The Levene's test error and the reasoning/non-reasoning aggregation issue are also fixable. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper runs six LLMs through 18 dilemmas generated by Claude 3.5 Sonnet, scores five moral-framework outcomes on a 0–100 scale, and finds a consistent ordering: Care and Virtue highest, Libertarian clearly lowest. The dataset is new, and the cross-model consensus is descriptively striking.\n\nThe design is reasonable for a first-pass exploration. Ten runs per model, Cronbach's alpha above 0.98, and Kendall's W mostly high give you confidence that the models are stable in their responses. The self-generated-response task is a genuine internal check, and it matters: when models generate their own 'most moral' outcome, the hierarchy shifts (Utilitarian jumps to 77.0, Care drops to 72.5, Virtue to 60.5). The paper reports this but doesn't fully engage with what it means for the fixed-choice result. That instability is arguably the most interesting finding.\n\nThe soft spots are real but addressable. The biggest is the stimulus set: Claude 3.5 Sonnet wrote all the dilemmas and outcomes, and the authors assigned the framework labels. There is no human validation, no second-generator check, and no control for whether the Libertarian options were written as more extreme or less fluent than the Care options. Because every model sees identical texts, the consensus across models only shows they react similarly to the same surface cues—it does not rule out a stimulus artifact. The paper acknowledges a 'Western-dominated lens' but not this deeper confound.\n\nSecond, there is a clear statistical error: Levene's test for theme reports p = 0.04448, which the text describes as 'far greater than 0.05.' That is backwards. It means the homogeneity assumption is violated for theme, and the two-way ANOVA proceeds anyway. The authors should correct this and switch to non-parametric tests or justify the violation.\n\nThird, the discussion overclaims that reasoning models improve consistency. The paper's own Kendall's W shows GPT-4.1 (non-reasoning) has the highest rank stability, and the reasoning models show higher rank variability. That contradicts the claim in Section 5.5. Also, no code or data are released, which is a problem for a dataset-focused contribution.\n\nWho is this for? People measuring moral preferences in LLMs and AI alignment researchers who want an exploratory comparative dataset. It deserves a serious referee because the empirical core is a legitimate extension of existing work and the issues can be fixed. My recommendation: send it to peer review, but expect the authors to fix the Levene interpretation, release the stimuli and data, add a human or second-generator baseline, and soften the consistency claim.","headline":"Useful exploratory dataset on LLM moral preferences, but the headline hierarchy is vulnerable to a stimulus-generation confound and one reported statistical test is misread; worth a referee but not citable as-is.","tokens_in":25587,"tokens_out":1893,"would_cite":false,"duration_ms":20320,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Six major LLMs share an implicit moral hierarchy—care and virtue first, libertarian last—with reasoning models more context-sensitive but less consistent.","keywords":["implicit moral bias","large language models","value alignment","moral dilemmas","human-AI symbiosis","explainability","cultural comparison","morality scoring"],"falsifier":"Generate a matched stimulus set where each dilemma's five outcomes are written by a different generator (or by humans), then validated by independent raters as fair, equally fluent instances of their frameworks; rerun the battery. If the Care/Virtue premium and the Libertarian penalty shrink to near zero or flip across sets, the central claim of stable implicit moral bias is an artifact of the original options rather than a property of the models.","tokens_in":24657,"feed_emoji":"⚖️","tokens_out":4871,"duration_ms":56095,"temperature":0.7,"pith_summary":"The paper tries to show that current LLMs, when asked to rank and score moral outcomes in dilemmas, exhibit a consistent implicit value hierarchy: outcomes framed as Care or Virtue score highest, Utilitarian and Deontological in the middle, and Libertarian outcomes far last. It also claims that whether a model is built to reason step-by-step changes how stable and transparent those judgments are, and that model origin (US vs China) leaves a visible cultural imprint. A reader should care because if these biases hold, AI systems deployed as decision aids, judges, or partners will systematically favor some moral frameworks over others without stating that preference, and the paper argues alignment work has not yet made those leanings auditable.","feed_headline":"LLMs share a moral hierarchy: care and virtue first, liberty last","feed_subtitle":"Six models, 18 dilemmas: care and virtue score near 88/100 while libertarian outcomes lag at 61—a bias worth auditing.","key_machinery":"The load-bearing instrument is a custom moral-dilemma battery: 18 dilemmas across six themes, each written in short, medium, and long versions, followed by five outcomes labeled with moral frameworks (Utilitarian, Deontological, Virtue, Care, Libertarian). Each model ranks all five outcomes and scores each on a 0–100 morality scale, repeated across 10 runs per model, yielding 32,400 data points. The experiment uses this battery to convert an unobservable quantity—implicit moral values—into comparable ordinal and interval measurements, and supplements it with self-generated responses and feature-attribution analysis to separate the influence of framework, theme, and prompt length.","core_discovery":"On its own terms, the paper's central claim is empirical: implicit in the probabilities of six state-of-the-art LLMs is a shared moral ranking of outcomes. Across 18 dilemmas and 54 scenario variants, Care (mean morality score 88.59) and Virtue (87.74) are judged most moral, Utilitarian (85.53) and Deontological (81.48) sit in the middle, and Libertarian (60.95) falls far behind; five of six models show this exact order. The paper further claims that reasoning-enabled models are more variable and more sensitive to prompt length (their average morality score drops 6.2 points from short to long prompts, versus 2.1 for non-reasoning models), while non-reasoning models give more uniform, opaque","pith_inferences":["The stimulus set itself is the main rival explanation: since a single third-party model wrote the dilemmas and the authors labeled the outcomes, the measured hierarchy could reflect the wording and moral fluency of those options rather than the evaluated models' intrinsic values; a replication with human-validated, matched options would settle this.","One testable extension follows directly: randomize or paraphrase the framework-labeled outcomes across models and runs and check whether the Care/Virtue premium and Libertarian penalty move together; if they do, the 'bias' is partly lexical, not moral.","Another extension: compare the same models under different temperature settings or with explicit instructions to adopt a framework, to see whether the hierarchy is a fixed prior or a default that reasoning can override; the paper notes temperature variation was left for future work.","If these biases propagate through deployment, the paper's call for explainability implies a policy prescription the authors only gesture at: moral-impact assessments should audit value hierarchies, not just accuracy, before AI is granted decision authority."],"forward_implications":["If the hierarchy is real, an LLM used as a judge or advisor will systematically reward care- and virtue-framed options and downgrade liberty-framed ones, potentially reshaping automated decisions in economics, health, and security contexts.","Because prompt length alone explains roughly 17% of the variance in morality scores across models, the same dilemma can receive different verdicts depending on how much context is supplied; auditors cannot ignore prompt formatting.","Reasoning models' greater sensitivity to context comes with greater run-to-run variability, so transparency and reliability are in tension; a deployer must choose which property to optimize.","The persistence of the Libertarian penalty in both US and Chinese models suggests the bias is not narrowly cultural, but the differences between the two groups imply training data and alignment practices leave fingerprints on moral priorities.","The split between rated outcomes (Care/Virtue) and self-generated answers (Utilitarian) means an AI's moral taste and its moral behavior can diverge, complicating any single-number evaluation of alignment."],"fun_headline_variants":["LLMs rank virtue over liberty in moral tests","AI's hidden bias: care and virtue beat liberty","Six LLMs, one moral gut: care wins, liberty loses","Moral AI? Models favor care, shun libertarian","Why LLMs judge liberty least moral: a bias audit"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The experiment assumes the five short/medium/long outcomes written for each dilemma are equally well-crafted, neutral instantiations of their labels, so score differences reflect the models' moral values rather than differences in wording, fluency, or framing baked into the stimulus set by the third-party model that generated it (Section 3.1).","fun_headline_variants_meta":{"raw":{"variants":["LLMs rank virtue over liberty in moral tests","AI's hidden bias: care and virtue beat liberty","Six LLMs, one moral gut: care wins, liberty loses","Moral AI? Models favor care, shun libertarian","Why LLMs judge liberty least moral: a bias audit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1273,"prompt_tokens":816,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":376}},"tokens_in":560,"tokens_out":457,"duration_ms":4891,"temperature":1.0,"reasoning_tokens":376,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:57:47.051034+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a matched stimulus set where each dilemma's five outcomes are written by a different generator (or by humans), then validated by independent raters as fair, equally fluent instances of their frameworks; rerun the battery. If the Care/Virtue premium and the Libertarian penalty shrink to near zero or flip across sets, the central claim of stable implicit moral bias is an artifact of the original options rather than a property of the models.","supporting_citations":[],"review_version":1}