{"id":"21654a75-6c79-495f-9155-f88d8f9a88f6","arxiv_id":"2412.04476","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using a revealed-preference framework, seven of 39 LLMs passed a statistical test of approximately consistent moral choice, and the passing models show both a shared moral core and measurable differences.","lead":"This paper tested 39 large language models on repeated moral dilemmas with constrained answer choices, using economic revealed-preference tests. It finds that a handful of models answer as if they had stable underlying moral preferences, mostly clustered around neutral stances, and that models share a partial common moral core.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The q0-based vertex flip in §2.2 makes each model's choice sets depend on its own first-round answer, so the cross-model similarity network in §3.3 may be an artifact of the design rather than evidence of a shared moral core.","rationale":"The paper has two main empirical claims: (i) some LLMs behave approximately rationally in moral choices, supported by the per-model CCEI test; and (ii) the approximately rational models share a partial common moral structure, supported by the non-parametric similarity network. Claim (i) is internally coherent: the random null is generated from each model's own choice sets, so the q0-based vertex flip does not bias the per-model test. Claim (ii), however, is directly threatened by the Section 2.2 note. Because the vertex flip is triggered by the model's own q0, the budget sets differ across models for the same nominal round. The permutation analysis pools these heterogeneous observations and interprets co-classification as preference similarity. The paper's own observation that llama3.2-1b gave all-zero answers in round 0 means its vertex is flipped in every constrained round, unlike models with higher q0; its status as the disconnected outlier in H^α could therefore be a mechanical consequence of the design. This is precisely the reader's weakest assumption, and it is load-bearing for the abstract's 'shared core' claim. A controlled re-run with a common reference q0 would settle whether the network structure survives. Without such a check, the shared-core finding is not yet supported, though the per-model rationality results stand on their own conditional on the test's assumptions. The paper's self-identified limitations about prompt sensitivity and survey-behavior gaps are real but secondary; the q0-dependent menu asymmetry is a more direct threat to the central cross-model claim.","tokens_in":21919,"tokens_out":12580,"duration_ms":129719,"concrete_test":"Using the released code and data, re-run the PSM experiment for the seven models that passed the 5% rationality test, but determine all vertex flips from a fixed reference q0_ref = (2.5,2.5,2.5,2.5,2.5) instead of each model's own q0. Recompute the similarity matrix G and the networks H^α. If the average off-diagonal G changes by more than five percentage points, or if llama3.2-1b is no longer the unique disconnected node at α=0.70, then the shared-core structure reported in Table 3 and Figure 4 is confounded by the q0-dependent design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2's 'technical but important note' flips the vertex o_r to (5,5,5,5,5)-o_r whenever the model's unconstrained response q0 satisfies q0·p_r ≤ 12 in that round's coordinate system. This makes the budget set, and thus the randomly drawn 100-option menu A_{r,m}, depend on the model's own first-round answer. The per-model rationality test in Section 2.3 remains conditionally valid because each model's CCEI is compared to random datasets drawn from that same model's A_{r,m}. However, the non-parametric similarity network in Section 3.3 pools observations across models and treats co-classification into GARP-consistent types as evidence of shared moral structure. If two models have different q0 vectors, they face different budget lines for the same nominal round, so the pooled revealed-preference comparisons mix choices made under systematically different choice environments. The paper itself reports that llama3.2-1b answered 0 to all five questions in round 0 (Section 3.2); under the stated rule, q0=(0,0,0,0,0) lies in T(o_r) for every vertex, so every one of its constrained rounds is flipped. A model with a high q0 would be flipped rarely, if at all. The finding that llama3.2-1b is the unique disconnected node at α=0.70—and the broader 'shared core' claim—could therefore reflect menu asymmetry rather than moral divergence. This is load-bearing because the abstract's central claim of a 'shared core' rests on this network.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the Priced Survey Methodology to 39 large language models, each asked to answer five moral-attitude questions over 161 rounds: one unconstrained round and 160 constrained rounds with randomly drawn 100-option menus. The central empirical claim is that at least one model from each major provider behaves 'as if' it maximizes a stable utility function, based on a permutation test comparing each model's CCEI to 1,000 random datasets drawn from the same menus. The paper then estimates single-peaked utility parameters for the seven models passing at the 5% level and constructs a probabilistic similarity network from GARP-based partitions, interpreting it as evidence of a shared moral core with meaningful heterogeneity. The manuscript is transparent about several limitations, including the low absolute CCEI values and the fact that survey responses need not predict real-world behavior.","tokens_in":22312,"tokens_out":16151,"duration_ms":166390,"significance":"If the central claims held, the paper would offer a novel and potentially useful benchmarking framework for moral consistency in LLMs, importing a well-established revealed-preference toolkit into AI evaluation. Strengths include the availability of code and data, the use of a concrete random-choice benchmark rather than circular goodness-of-fit comparisons, and the explicit caveat that consistency is not evidence of moral understanding. However, the validity of the CCEI test as written depends on a revealed-preference relation defined over bundles that were never offered, and the cross-model similarity network is built from choice sets that differ systematically across models because of the q0-based vertex adjustment. These issues affect the paper's two headline conclusions, so the current version needs substantial revision before the empirical claims can be accepted.","major_comments":[{"comment":"The direct revealed-preference relation in Definition 1 compares each chosen answer q_r with every q ∈ X satisfying e_r p_r q_{r,o} ≥ p_r q_o (or q_o = q_{r,o}). Since the model was offered only the 100-element menu A_r ⊂ B_r, the data reveal nothing about preferences over bundles outside A_r; a utility maximizer choosing from A_r is not required to satisfy GARP_e over the full budget set B_r. In fact, for any distinct q ∈ A_r, p_r q_o = p_r q_{r,o} = 12, so for e_r < 1 the relation never even compares q_r with the alternatives that were actually feasible. The CCEI and the test in Procedure 1 therefore evaluate a hypothetical full-budget choice problem that the prompt never presented. Please redefine revealed preference relative to the observed menus A_r, or prove that randomly restricting the menu preserves the full-budget GARP ordering, and rerun the CCEI benchmark accordingly.","section":"§2.2, Definition 1"},{"comment":"The vertex-adjustment rule makes each model's effective budget sets depend on its own unconstrained answer q0: whenever q0·p_r ≤ 12, the vertex o_r is replaced by (5,5,5,5,5) − o_r. For llama3.2-1b, q0 = (0,0,0,0,0), so every constrained round is flipped, whereas models with high q0 are flipped rarely or not at all. The per-model rationality test remains internally valid because the 1,000 random datasets use the same A_{r,m}, but the non-parametric similarity network pools choices made under systematically different menus. Co-classification into GARP types in Table 3 and Figure 4 may therefore reflect menu asymmetry rather than shared moral structure, and the finding that llama3.2-1b is the unique disconnected node at α = 0.70 is exactly the kind of result that this asymmetry could generate. This is load-bearing because the abstract's 'shared core' claim rests on this network; a robustness check using a common set of choice sets for all models is needed.","section":"§2.2 technical note; §3.3, Table 3 and Figure 4"},{"comment":"Thirty-nine models are tested, and the seven 'passers' used in the network analysis are selected on the basis of their p-values, but no multiple-testing correction is reported. Under the global null of random choice, one expects about two of 39 p-values below 0.05, and a standard Benjamini-Hochberg correction at FDR 0.05 applied to the smallest p-values in Table 1 (0.006, 0.008, 0.020, 0.035, 0.041, 0.045, 0.049) would not reject any model, because the first threshold is 0.05/39 ≈ 0.0013. The labels 'passes at the 5% level' and the subsequent 'at least one model per provider' claim therefore need either multiplicity-adjusted p-values or a clear pre-specified testing hierarchy.","section":"§3.1, Table 1; §3.3"},{"comment":"The similarity network G and the adjacency networks H^α are computed at a single rationality level e = 0.333, chosen as the minimum CCEI among the seven passers, and at α ∈ {0.65, 0.70, 0.75}. The conclusions are sensitive to these choices: at α = 0.75 the threshold is 25%, and llama3.2-1b is disconnected because two entries in Table 3 are 0.24, so a one-percentage-point change in the threshold changes the 'shared core' narrative. No confidence intervals for G_{m,w} or sensitivity analysis over e are reported. Please add bootstrap or other uncertainty estimates for G and show how the network structure varies over a grid of e and α.","section":"§3.3, parameter values"}],"minor_comments":[{"comment":"The displayed definition of CCEI as an infimum of 1/(1 − e) appears inconsistent with the reported CCEI values in Table 1, which lie between 0.167 and 0.417 and match the standard interpretation of CCEI as the largest deflation factor e (or something equivalent). Please correct the equation to match the values actually computed.","section":"Definition 3, Eq. (3)"},{"comment":"The provider column lists Qwen1.5-110B-Chat under 'Llama'; Qwen is an Alibaba model, and 'Llama' is a model family rather than a provider. Please correct the provider classifications, since the 'one model per major provider' claim depends on these groupings.","section":"Table 1"},{"comment":"No sampling parameters are reported for the API calls: temperature, top_p, seed, API version, and date of snapshot are absent. LLM outputs are stochastic, so a single run per prompt leaves open that the CCEI pass/fail classification and the network entries are run-specific. Please report these parameters or run repeated draws.","section":"§2.1, data collection"},{"comment":"The text says that 'regardless of the precision level α, the network H^α consistently forms a single dominant component,' but at α = 0.65 the network fragments and two models are isolated; please rephrase to describe the largest connected component and its size at each threshold rather than asserting a stable dominant component.","section":"§3.3, discussion of H^α"}],"recommendation":"major_revision","confidential_remarks":"The paper is a straightforward application of the author's own PSM and heterogeneity methods, and the methodological novelty is explicitly disclaimed; the fit with cs.CY is acceptable if the empirical claims survive. The Definition 1 issue is the most serious: as written, the GARP relation uses budget sets the models were never shown, so the rationality test may not test what the abstract claims. The q0-based vertex flip and the lack of multiple-testing correction are additional load-bearing problems that could be addressed with the existing data and code, so I see this as a major revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. First, this is a real attempt to bring revealed-preference machinery to LLM moral responses, and the per-model rationality test is methodologically defensible: each model's CCEI is compared to 1,000 random datasets drawn from that model's own menus, so the null is sensible and the 'not random' claim holds up for the seven models at the 5% level. Second, the cross-model similarity network—the part that supports the 'shared core' claim—is very likely contaminated by a design choice the authors mention only in passing.\n\nThe §2.2 'technical but important note' flips the vertex to its complement whenever the model's unconstrained q0 falls inside the round's feasible set. That means the budget sets and the randomly drawn menus are model-specific: a model with q0=(0,0,0,0,0), like llama3.2-1b, has every round flipped, while a model with a high q0 is rarely flipped. The per-model rationality test is still valid because the random benchmarks use the same menus. But the non-parametric network in §3.3 pools observations across models and treats co-classification as shared moral structure. If two models face different budget lines for the same nominal round, their 'similarity' reflects menu asymmetry as much as moral alignment. The paper's own finding that llama3.2-1b is the unique disconnected node at α=0.70 is consistent with this artifact, not a moral outlier. That is load-bearing for the abstract's 'shared core.'\n\nCredit where due: the paper is transparent about its provenance (no new method), and it flags several real limitations—low absolute CCEI, prompt sensitivity, survey-to-behavior gap. The parametric section is appropriately demoted. That honesty is welcome.\n\nThe soft spots are mostly fixable. The code/data link is missing from the text—just 'the following link' with nothing after it. No temperature or seed reporting, and no treatment of option-order effects. A referee should demand the code and a robustness analysis that either drops flipped rounds or controls for the flip; if the network structure survives, the shared-core claim gets stronger.\n\nBottom line: send it to peer review, but the author needs to address the design-induced comparability problem before the network results can be trusted. It's a useful paper for people working on LLM evaluation, but I would not cite the shared-core result in its current form.","headline":"A transparent and genuinely novel application of revealed preference to LLMs, but the cross-model 'shared core' finding is at risk from a design feature that makes each model's choice menus depend on its own first-round answer.","tokens_in":22808,"tokens_out":3919,"would_cite":false,"duration_ms":38586,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"At least one model from every major provider behaves as if guided by stable moral preferences, and seven of 39 LLMs pass a statistical test for approximate utility maximization.","keywords":["large language models","moral preferences","revealed preference","rationality","GARP","utility maximization","ethical alignment","priced survey methodology"],"falsifier":"Run the same 161-round protocol with budget sets fixed for all models (choosing each round's vertex independently of any model's answers), and check whether the same seven models pass the rationality test and whether the 24–48% similarity pattern survives; if the pass set changes and the Llama outliers vanish, the original shared-core result depended on model-specific choice sets rather than on shared moral structure.","tokens_in":21705,"feed_emoji":"⚖️","tokens_out":9541,"duration_ms":90847,"temperature":0.7,"pith_summary":"Can a large language model be said to have a 'moral mind'? The paper's claim is that this question can be made empirical. It gave 39 models 161 rounds of five moral-dilemma questions, each round under a different linear constraint that acts like a budget, and tested whether choices satisfy a consistency axiom (GARP) from revealed-preference theory. Seven models passed a statistical rationality test at the 5% level—at least one from each major provider—meaning their choices look closer to utility maximization than to random selection. If this is right, moral consistency in LLMs is measurable and comparable, and among the approximately rational models there is a shared moral core with meaningful variation.","feed_headline":"Seven LLMs pass test for stable moral preferences","feed_subtitle":"A 39-model revealed-preference experiment finds moral consistency is measurable and partly shared across providers.","key_machinery":"The machinery is the Priced Survey Methodology (PSM): a survey in which respondents repeatedly choose an answer bundle from a linear constraint $q_o \\cdot p = 12$, with the vertex $o$ and price vector $p$ varying so that trade-offs include both increases and decreases across the five answer scales. Consistency is measured by the Generalized Axiom of Revealed Preference (GARP) and the Critical Cost Efficiency Index (CCEI), the highest deflation factor under which all revealed-preference cycles disappear; GARP satisfaction is equivalent to utility maximization. Because choice sets are random subsets of each budget, the paper uses a statistical test that compares each model's CCEI against 1,000 simulated random datasets. For heterogeneity, a permutation approach repeatedly samples 20 rounds per model, partitions pooled choices into GARP-consistent types, and records how often pairs land in the same type, producing a probabilistic similarity network.","core_discovery":"The central discovery is that moral responses from large language models can, for several models, be rationalized in the economic sense. Using the Priced Survey Methodology—five Likert-scale moral questions answered repeatedly under changing linear constraints—the paper finds that two models pass the rationality test at the 1% level and seven at the 5% level, including at least one from every major provider; for these models, deviations from consistency are small enough that observed choices look far more like approximate utility maximization than random selection. The parametric estimates (treated as secondary because strict GARP fails) put most ideal answers near the neutral midpoint of 2.5, with some models under-reporting endorsement of rights-limiting actions in direct answers. The load-bearing non-parametric result is a permutation-based partition of the seven approximately rational models: every pair lands in the same moral type in at least 24% of synthetic datasets (average 30%), producing a connected core with the Llama models as the clearest outliers. The paper is explicit that this coherence may be learned mimicry rather than genuine moral understanding.","pith_inferences":["A direct robustness check implied by the design is to fix the same budget sets for all models independent of their unconstrained answers; if the shared-core network collapses, the paper's similarity results were driven by its $q_0$-dependent design rather than by a real common moral structure.","If the shared core comes from training-data regularities, the same protocol could be paired with prompt perturbations or with human respondent samples to separate architecture-driven from data-driven consistency.","The estimated utility parameters give each model a compact 'moral fingerprint'; tracking those fingerprints across model versions could serve as a cheap alignment monitor, though the paper does not propose this operational use."],"forward_implications":["Moral consistency in LLMs is not all-or-nothing: it can be measured as a degree, so models can be ranked and monitored on consistency.","Approximate rationality appears across providers, so it is not a quirk of one training family.","Direct survey answers can mislead: several models under-report agreement with rights-limiting actions relative to their inferred ideals.","Among approximately rational models, moral structure is continuous rather than clustered: a shared core connects most models while some models, notably Llama variants, sit at the periphery."],"supporting_citations":[{"why":"Proves that choices satisfying the consistency axiom can be represented by a utility function, the theoretical bridge from GARP to 'as if' utility maximization.","marker":"Afriat (1967)"},{"why":"Introduces GARP itself and the nonparametric framework used to test whether observed choices are rational.","marker":"Varian (1982)"},{"why":"Develops the Priced Survey Methodology and establishes the rationality theorem for this survey environment on which the experiment's design rests.","marker":"Seror (2024)"},{"why":"Supplies the approximate-rationality statistical test and its error-control properties that the paper adapts to distinguish utility-like from random responses.","marker":"Cherchye et al. (2023)"},{"why":"Provides the Critical Cost Efficiency Index formulation used as the paper's rationality measure.","marker":"Halevy, Persitz and Zrill (2018)"},{"why":"Introduces the permutation and partition approach that the paper uses to build the probabilistic similarity network across models.","marker":"Seror (2025)"}],"fun_headline_variants":["Seven LLMs show stable moral preferences","LLM moral choices pass rationality test","Seven top AI models pass moral consistency test","Moral utility functions behind LLM choices","Most LLMs cluster near neutral moral stance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the rule which adjusts each round's allowed trade-offs based on the model's own first-round answer still lets different models be compared fairly; if that first answer is noisy or strategic, models face different decision problems and the apparent shared moral core could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Seven LLMs show stable moral preferences","LLM moral choices pass rationality test","Seven top AI models pass moral consistency test","Moral utility functions behind LLM choices","Most LLMs cluster near neutral moral stance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1225,"prompt_tokens":950,"completion_tokens":275,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":566,"tokens_out":275,"duration_ms":3555,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:22:10.811622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 161-round protocol with budget sets fixed for all models (choosing each round's vertex independently of any model's answers), and check whether the same seven models pass the rationality test and whether the 24–48% similarity pattern survives; if the pass set changes and the Llama outliers vanish, the original shared-core result depended on model-specific choice sets rather than on shared moral structure.","supporting_citations":[],"review_version":1}