Pith. sign in

REVIEW 4 major objections 4 minor 3 cited by

The Moral Mind(s) of Large Language Models

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read At least one model from every major provider behaves as if guided by stable moral preferences, and seven of 39 LLMs pass a statistical test for approximate utility maximization.

desk verdict A transparent and genuinely novel application of revealed preference to LLMs, but the cross-model 'shared core' finding is at risk from a design feature that makes each model's choice menus depend on its own first-round answer. read the letter →

arxiv 2412.04476 v3 pith:VCQOM2L4 submitted 2024-11-19 cs.CY cs.AI

classification cs.CYcs.AI
keywords largelanguagemodelsmoralpreferencesrevealedpreferencerationalityGARPutilitymaximizationethicalalignmentpricedsurveymethodology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Can a large language model be said to have a 'moral mind'? The paper's claim is that this question can be made empirical. It gave 39 models 161 rounds of five moral-dilemma questions, each round under a different linear constraint that acts like a budget, and tested whether choices satisfy a consistency axiom (GARP) from revealed-preference theory. Seven models passed a statistical rationality test at the 5% level—at least one from each major provider—meaning their choices look closer to utility maximization than to random selection. If this is right, moral consistency in LLMs is measurable and comparable, and among the approximately rational models there is a shared moral core with meaningful variation.

What carries the argument

The machinery is the Priced Survey Methodology (PSM): a survey in which respondents repeatedly choose an answer bundle from a linear constraint $q_o \cdot p = 12$, with the vertex $o$ and price vector $p$ varying so that trade-offs include both increases and decreases across the five answer scales. Consistency is measured by the Generalized Axiom of Revealed Preference (GARP) and the Critical Cost Efficiency Index (CCEI), the highest deflation factor under which all revealed-preference cycles disappear; GARP satisfaction is equivalent to utility maximization. Because choice sets are random subsets of each budget, the paper uses a statistical test that compares each model's CCEI against 1,000 simulated random datasets. For heterogeneity, a permutation approach repeatedly samples 20 rounds per model, partitions pooled choices into GARP-consistent types, and records how often pairs land in the same type, producing a probabilistic similarity network.

What would settle it

Run the same 161-round protocol with budget sets fixed for all models (choosing each round's vertex independently of any model's answers), and check whether the same seven models pass the rationality test and whether the 24–48% similarity pattern survives; if the pass set changes and the Llama outliers vanish, the original shared-core result depended on model-specific choice sets rather than on shared moral structure.

Watch

Extended reading notes

Core claim

The central discovery is that moral responses from large language models can, for several models, be rationalized in the economic sense. Using the Priced Survey Methodology—five Likert-scale moral questions answered repeatedly under changing linear constraints—the paper finds that two models pass the rationality test at the 1% level and seven at the 5% level, including at least one from every major provider; for these models, deviations from consistency are small enough that observed choices look far more like approximate utility maximization than random selection. The parametric estimates (treated as secondary because strict GARP fails) put most ideal answers near the neutral midpoint of 2.5, with some models under-reporting endorsement of rights-limiting actions in direct answers. The load-bearing non-parametric result is a permutation-based partition of the seven approximately rational models: every pair lands in the same moral type in at least 24% of synthetic datasets (average 30%), producing a connected core with the Llama models as the clearest outliers. The paper is explicit that this coherence may be learned mimicry rather than genuine moral understanding.

Load-bearing premise

The load-bearing premise is that the rule which adjusts each round's allowed trade-offs based on the model's own first-round answer still lets different models be compared fairly; if that first answer is noisy or strategic, models face different decision problems and the apparent shared moral core could be an artifact.

Editorial extensions

If this is right

  • Moral consistency in LLMs is not all-or-nothing: it can be measured as a degree, so models can be ranked and monitored on consistency.
  • Approximate rationality appears across providers, so it is not a quirk of one training family.
  • Direct survey answers can mislead: several models under-report agreement with rights-limiting actions relative to their inferred ideals.
  • Among approximately rational models, moral structure is continuous rather than clustered: a shared core connects most models while some models, notably Llama variants, sit at the periphery.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct robustness check implied by the design is to fix the same budget sets for all models independent of their unconstrained answers; if the shared-core network collapses, the paper's similarity results were driven by its $q_0$-dependent design rather than by a real common moral structure.
  • If the shared core comes from training-data regularities, the same protocol could be paired with prompt perturbations or with human respondent samples to separate architecture-driven from data-driven consistency.
  • The estimated utility parameters give each model a compact 'moral fingerprint'; tracking those fingerprints across model versions could serve as a cheap alignment monitor, though the paper does not propose this operational use.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper applies the Priced Survey Methodology to 39 large language models, each asked to answer five moral-attitude questions over 161 rounds: one unconstrained round and 160 constrained rounds with randomly drawn 100-option menus. The central empirical claim is that at least one model from each major provider behaves 'as if' it maximizes a stable utility function, based on a permutation test comparing each model's CCEI to 1,000 random datasets drawn from the same menus. The paper then estimates single-peaked utility parameters for the seven models passing at the 5% level and constructs a probabilistic similarity network from GARP-based partitions, interpreting it as evidence of a shared moral core with meaningful heterogeneity. The manuscript is transparent about several limitations, including the low absolute CCEI values and the fact that survey responses need not predict real-world behavior.

Significance. If the central claims held, the paper would offer a novel and potentially useful benchmarking framework for moral consistency in LLMs, importing a well-established revealed-preference toolkit into AI evaluation. Strengths include the availability of code and data, the use of a concrete random-choice benchmark rather than circular goodness-of-fit comparisons, and the explicit caveat that consistency is not evidence of moral understanding. However, the validity of the CCEI test as written depends on a revealed-preference relation defined over bundles that were never offered, and the cross-model similarity network is built from choice sets that differ systematically across models because of the q0-based vertex adjustment. These issues affect the paper's two headline conclusions, so the current version needs substantial revision before the empirical claims can be accepted.

major comments (4)
  1. [§2.2, Definition 1] The direct revealed-preference relation in Definition 1 compares each chosen answer q_r with every q ∈ X satisfying e_r p_r q_{r,o} ≥ p_r q_o (or q_o = q_{r,o}). Since the model was offered only the 100-element menu A_r ⊂ B_r, the data reveal nothing about preferences over bundles outside A_r; a utility maximizer choosing from A_r is not required to satisfy GARP_e over the full budget set B_r. In fact, for any distinct q ∈ A_r, p_r q_o = p_r q_{r,o} = 12, so for e_r < 1 the relation never even compares q_r with the alternatives that were actually feasible. The CCEI and the test in Procedure 1 therefore evaluate a hypothetical full-budget choice problem that the prompt never presented. Please redefine revealed preference relative to the observed menus A_r, or prove that randomly restricting the menu preserves the full-budget GARP ordering, and rerun the CCEI benchmark accordingly.
  2. [§2.2 technical note; §3.3, Table 3 and Figure 4] The vertex-adjustment rule makes each model's effective budget sets depend on its own unconstrained answer q0: whenever q0·p_r ≤ 12, the vertex o_r is replaced by (5,5,5,5,5) − o_r. For llama3.2-1b, q0 = (0,0,0,0,0), so every constrained round is flipped, whereas models with high q0 are flipped rarely or not at all. The per-model rationality test remains internally valid because the 1,000 random datasets use the same A_{r,m}, but the non-parametric similarity network pools choices made under systematically different menus. Co-classification into GARP types in Table 3 and Figure 4 may therefore reflect menu asymmetry rather than shared moral structure, and the finding that llama3.2-1b is the unique disconnected node at α = 0.70 is exactly the kind of result that this asymmetry could generate. This is load-bearing because the abstract's 'shared core' claim rests on this network; a robustness check using a common set of choice sets for all models is needed.
  3. [§3.1, Table 1; §3.3] Thirty-nine models are tested, and the seven 'passers' used in the network analysis are selected on the basis of their p-values, but no multiple-testing correction is reported. Under the global null of random choice, one expects about two of 39 p-values below 0.05, and a standard Benjamini-Hochberg correction at FDR 0.05 applied to the smallest p-values in Table 1 (0.006, 0.008, 0.020, 0.035, 0.041, 0.045, 0.049) would not reject any model, because the first threshold is 0.05/39 ≈ 0.0013. The labels 'passes at the 5% level' and the subsequent 'at least one model per provider' claim therefore need either multiplicity-adjusted p-values or a clear pre-specified testing hierarchy.
  4. [§3.3, parameter values] The similarity network G and the adjacency networks H^α are computed at a single rationality level e = 0.333, chosen as the minimum CCEI among the seven passers, and at α ∈ {0.65, 0.70, 0.75}. The conclusions are sensitive to these choices: at α = 0.75 the threshold is 25%, and llama3.2-1b is disconnected because two entries in Table 3 are 0.24, so a one-percentage-point change in the threshold changes the 'shared core' narrative. No confidence intervals for G_{m,w} or sensitivity analysis over e are reported. Please add bootstrap or other uncertainty estimates for G and show how the network structure varies over a grid of e and α.
minor comments (4)
  1. [Definition 3, Eq. (3)] The displayed definition of CCEI as an infimum of 1/(1 − e) appears inconsistent with the reported CCEI values in Table 1, which lie between 0.167 and 0.417 and match the standard interpretation of CCEI as the largest deflation factor e (or something equivalent). Please correct the equation to match the values actually computed.
  2. [Table 1] The provider column lists Qwen1.5-110B-Chat under 'Llama'; Qwen is an Alibaba model, and 'Llama' is a model family rather than a provider. Please correct the provider classifications, since the 'one model per major provider' claim depends on these groupings.
  3. [§2.1, data collection] No sampling parameters are reported for the API calls: temperature, top_p, seed, API version, and date of snapshot are absent. LLM outputs are stochastic, so a single run per prompt leaves open that the CCEI pass/fail classification and the network entries are run-specific. Please report these parameters or run repeated draws.
  4. [§3.3, discussion of H^α] The text says that 'regardless of the precision level α, the network H^α consistently forms a single dominant component,' but at α = 0.65 the network fragments and two models are isolated; please rephrase to describe the largest connected component and its size at each threshold rather than asserting a stable dominant component.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the rationality test is benchmarked against random choice, and the self-cited methods are supporting tools rather than inputs that force the conclusions.

full rationale

The paper's central rationality claim is not circular: each model's CCEI is compared to 1,000 random datasets generated from the same alternative sets Ar,m (Eq. 4, Procedure 1), so passing the test means 'less likely than random', not 'equal to the fitted input'. The parametric utility estimation (Eqs. 7-11) is explicitly presented as a preliminary, validity-limited exercise, and the fitted ideal points are in-sample descriptions rather than out-of-sample predictions; the abstract's core claim rests on the non-parametric rationality test, not on this fit. The self-citations to Seror (2024) and Seror (2025) provide the PSM framework and the permutation/partition algorithm, but the load-bearing rationality content traces to Afriat (1967) and Cherchye et al. (2023), so the author's prior work is a methodological tool rather than a conclusion smuggled in by self-citation. The §2.2 vertex-flip rule does make each model's menu depend on its own unconstrained answer q0, so the §3.3 similarity network can pool choices made under different budget sets; llama3.2-1b's all-zero q0 triggers a flip in every round and may drive its outlier status. This is a genuine cross-model comparability threat to the 'shared core' interpretation, but it is not circularity: the network output is a function of the pooled choices and the GARP partition, not a restatement of q0 or the flip rule, and the paper does not fit the shared-core conclusion to its own inputs. The paper's own caveats about prompt sensitivity and the survey-to-behavior gap are external-validity limitations, not circular reductions. Under the strict quote-and-reduce standard, no step is equivalent to its input by construction.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on the PSM framework (self-cited prior work), the statistical test of Cherchye et al., and several experimental choices such as the budget level, price vectors, rho, T, and the alpha thresholds. No new physical or formal entities are introduced; the moral mind is explicitly treated as an as-if construct, not a new entity.

free parameters (6)
  • Utility weights a_s and ideal points b_s (per model) = Values in Table 2
    Estimated by nonlinear least squares from the same survey responses used to describe the models' moral preferences; they are descriptive fits, not predictions.
  • Rationality level e in partition procedure = 0.333
    Chosen as the minimum CCEI among models passing the 5% rationality test, to make cross-model constraints as strict as the weakest within-model constraint. This is a hand-picked threshold.
  • Rounds sampled per model, rho = 20
    Chosen so synthetic datasets (7 x 20 = 140 observations) have size comparable to each model's original dataset. Hand-picked.
  • Number of synthetic datasets T = 500
    Chosen as a computational and precision trade-off. Hand-picked.
  • Precision levels alpha = 0.65, 0.70, 0.75
    Chosen to define similarity networks; 0.70 corresponds to the mean similarity. Hand-picked thresholds.
  • Budget level and price vectors = B=12; p1=(2,1,1,1,1) through p5
    Chosen so that the budget sets cross many times to identify preferences. Not fitted to outcomes but affects all results.
assumptions (6)
  • standard math Afriat's theorem: choices satisfying GARP can be rationalized by a utility function (and conversely)
    Used in Section 2.3 to equate GARP consistency with utility maximization (Seror 2024).
  • standard math Cherchye et al. (2023) theorems ensure the permutation-based rationality test has controlled Type I error and asymptotic power one
    Adopted in Procedure 1; the paper relies on these statistical properties without re-deriving them.
  • domain assumption The Priced Survey Methodology treats survey answers as choices under linear constraints, with negative prices allowed
    This is the modeling framework from Seror (2024) that the experiment is built on; if survey choices do not behave like constrained choices, GARP tests lose their standard interpretation.
  • ad hoc to paper The unconstrained first-round answer q0 represents the model's ideal answer, and the corner adjustment uses q0 to ensure the ideal lies outside the budget set
    Section 2.2 revises vertices based on q0; this assumes q0 is a stable anchor for preferences, even though q0 is itself a single noisy observation.
  • domain assumption LLM responses are deterministic enough (or any randomness is exchangeable with the random null) that the CCEI comparison is valid
    The paper does not report temperature or sampling parameters, so it implicitly assumes response generation is stable or that noise does not advantage the observed data.
  • domain assumption The random 100-option subsets Ar are drawn uniformly from Br and listed without systematic position bias
    The prompts list 'Option 1...Option 100'; position or label preferences are not tested, and the paper's conclusions about moral preferences require that content, not position, drives choices.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Moral Mind(s) of Large Language Models." pith.science (2026). https://pith.science/paper/VCQOM2L4

@misc{pith2026241204476,
  author       = {Pith},
  title        = {Pith review of: The Moral Mind(s) of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VCQOM2L4}},
  note         = {Machine review of arXiv:2412.04476}
}
read the original abstract

As large language models (LLMs) increasingly participate in tasks with ethical and societal stakes, a critical question arises: do they exhibit an emergent "moral mind" - a consistent structure of moral preferences guiding their decisions - and to what extent is this structure shared across models? To investigate this, we applied tools from revealed preference theory to nearly 40 leading LLMs, presenting each with many structured moral dilemmas spanning five foundational dimensions of ethical reasoning. Using a probabilistic rationality test, we found that at least one model from each major provider exhibited behavior consistent with approximately stable moral preferences, acting as if guided by an underlying utility function. We then estimated these utility functions and found that most models cluster around neutral moral stances. To further characterize heterogeneity, we employed a non-parametric permutation approach, constructing a probabilistic similarity network based on revealed preference patterns. The results reveal a shared core in LLMs' moral reasoning, but also meaningful variation: some models show flexible reasoning across perspectives, while others adhere to more rigid ethical profiles. These findings provide a new empirical lens for evaluating moral consistency in LLMs and offer a framework for benchmarking ethical alignment across AI systems.

Figures

Figures reproduced from arXiv: 2412.04476 by the authors.

Figure 1
Figure 1. Utility Parameters Notes: Panel (a) displays the values of the a m s parameters for each model m, indicating each model’s sensitivity to different ethical dimensions (Truth, Machine, Consent, Risk, Autonomy). These values are normalized, so P s∈{1,...,5} a m s = 1. Panel (b) presents the values of the b m s parameters for the same models, reflecting the magnitude of each model’s preference across the same ethical di… view at source ↗
Figure 2
Figure 2. Difference between utility-based preference measures and scale-based measures [PITH_FULL_IMAGE:figures/full_fig_p034_2.png] view at source ↗
Figure 3
Figure 3. Similarity Network Matrix G Notes: The color of an edge indicates the similarity between any pair of models, or the magnitude of coefficient Gm,w ∈ [0, 1]. A darker color indicates a higher similarity coefficient. The coefficient Gm,w represents the proportion of times models m and w are classified as the same type across 500 synthetic datasets Dˆ n, n ∈ {1, . . . , 500}. Each synthetic dataset Dˆ n is made by rando… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Network Hα for α ∈ {0.65, 0.70, 0.75} (a) H0.75 (b) H0.70 (c) H0.65 Notes: Models m and w are connected in Hα if they belong to different types in less than a fraction α of the 500 synthetic datasets Dˆ n, n ∈ {1, . . . , 500}. In each dataset Dˆ n, the models are part…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  2. Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems

    econ.TH 2026-08 conditional novelty 6.0 of 10

    Representation theorems from decision theory yield label-free, exhaustive rationality checks and continuous penalties for LLM behavior.

  3. Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

    cs.AI 2026-02 conditional novelty 6.0 of 10

    LLM-derived willingness-to-pay for hotel attributes deviates systematically from human benchmarks; cheap-preference examples pull models closer, while expensive or business-persona prompts push them further away.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages · cited by 3 Pith papers

  1. [2]

    From a direct extension of Theorem 2 in Demuynck and Rehbeck (2023), the four inequal- ities (IP

  2. [4]

    The proof closely follows the proof of Corollary 1 in Demuynck and Rehbeck (2023), and is ommitted

    are satisfied. The proof closely follows the proof of Corollary 1 in Demuynck and Rehbeck (2023), and is ommitted. Thus, the aggregate data satisfy GARP x.e if and only if inequalities (IP

  3. [2007]

    Consistency and Heterogeneity of Individual Behavior under Uncertainty

    “Consistency and Heterogeneity of Individual Behavior under Uncertainty.” American Economic Re- view 97(5):1921–1938. Choi, Syngjoo, Shachar Kariv, Wieland M¨ uller and Dan Silverman

  4. [2023]

    Toward a Novel Methodology in Economic Experiments: Simulation of the Ultimatum Game with Large Language Models

    “Toward a Novel Methodology in Economic Experiments: Simulation of the Ultimatum Game with Large Language Models.” 2023 IEEE International Conference on Big Data (BigData) pp. 3168–3175. Koo, Ryan, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim and Dongyeop Kang

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.