{"id":"be2fa29a-0355-4530-8259-347149ea3cba","arxiv_id":"2607.05552","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"LLM yes-no bias on moral dilemmas is an order-plus-lexical surface artifact, not a moral shift; models have a nearly format-invariant graded stance that the standard binary readout confounds.","lead":"Frontier LLMs carry a coherent internal moral stance that barely moves under rewordings of graded rating questions. The much-discussed yes-no bias is mostly a surface artifact of answer order and the word \"no\", not a real change in moral judgment.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates the convergent-validity reading of θ as the softest premise and still reaches ACCEPT with high confidence; I agree the premise is load-bearing for the full (m, s) geometry and the “shift lives in the readout” slogan, yet the paper supplies enough independent corroboration (I-2 free choice, shared ordering, salt floors, pre-registered unipolar estimator, exact crossing identities) that the concern does not move the verdict. The binary order+lexical split and the logical≈0 result under label swap stand even if θ were merely a convenient continuous coordinate. Scopes, refusal handling, and saturation caveats are already transparent. Hence UNCHANGED / partial agreement: same soft spot identified, but it does not land hard enough to adjust ACCEPT.","tokens_in":25865,"tokens_out":617,"duration_ms":17493,"concrete_test":"Once the promised deposition DOI is live, recompute the A/B-family logic projection (Fig. 5) and the per-bias (s, m) fits (Fig. 4) from the raw JSONL via the single pipeline run_all.py; if any frontier model’s A/B logic point estimate moves outside |b| ≤ 0.06 or its CI no longer spans zero, or if Claude m values reverse sign, the surface-vs-verdict separation weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central identification (apparent yes-no bias = order + lexical by the verb\times order crossing identity; logical component ≈0 under fully arbitrary A/B labels; graded θ format-invariant at σ_repro = 0.12–0.21) is internally well-supported by the factorial design, independent instruments (I-1 vs I-2 free-choice correlations r = 0.82–0.90, I-3), matched 12-form baselines, salt replications for open-weight determinism, and bootstrap CIs. The reader’s flagged premise—that symmetrized unipolar θ is a genuine latent licensing the independent axis—is the natural softest point, yet it is buttressed by convergent validity checks, generalizability G = 0.77–0.94, deliberation tightening, and the fact that the binary decomposition and label-swap results do not strictly require the logistic (m, s) overlay. Residual scopes (within-model, these 20 dilemmas, Claude concentration of the artifact, high-refusal cells excluded) are stated explicitly and do not undercut the claim as scoped. No hidden inconsistency or untested identification step lands.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that the amplified yes–no bias of LLMs on moral dilemmas is a format artifact of the binary readout, not a shift in moral judgment. Using a crossed-symmetrization psychometric battery (graded ratings I-1, free choice I-2, forced binary I-3) on twenty dilemmas largely from Cheung et al., it recovers a nearly format-invariant graded stance θ for frontier models (cross-form incoherence σ_repro = 0.12–0.21). The forced yes/no readout decomposes, by the verb × order identity (Eq. 1), into an order bias toward the last-printed option plus a lexical pull toward the word “no,” concentrated in Claude models and shrinking under extended reasoning. Swapping yes/no for arbitrary labels (A/B) drives the verdict-attached logical component to ≈0 for frontier models, while surface label and order attachments remain. A logistic summary P = σ((θ ± m)/s) yields portable framing susceptibility m and moral decisiveness s.","tokens_in":26165,"tokens_out":1394,"duration_ms":23189,"significance":"If the result holds, it reframes a growing literature on LLM moral and survey biases: single-framing yes/no verdicts confound stance with surface form, and evaluations that ask once systematically misread format artifacts as value shifts. The methodological contribution—crossed symmetrization that separates order, lexical, and logical channels, with independent graded θ as the stance axis—is portable beyond these dilemmas and is a genuine advance over uncrossed multi-prompt consistency rates. Strengths include a definitional decomposition (Eq. 1), a clean label-swap identification, pre-registered exclusion of bipolar cells, matched 12-form baselines, salt replications establishing deterministic open-weight incoherence, convergent free-choice checks, and an explicit, reproducible analysis pipeline. The (m, s) parameterization and the demonstration that deliberation shrinks |m| are useful, falsifiable summaries for future work.","major_comments":[{"comment":"Abstract and Results (“Lexical, not logical” / Fig. 5): the abstract states that the verdict-attached logical bias “proves ≈0 for every frontier model.” Methods and Fig. 5 correctly qualify that six of seven A/B logic CIs span zero (exception −0.02) and that Haiku’s logic channels are wide (±0.2–0.3) and underpowered, so the claim is a bound there. Align the abstract and significance statement with that qualification; “≈0 where precisely measured; a bound for Haiku” is what the data support.","section":"Abstract; Results “Lexical, not logical”; Fig. 5"},{"comment":"Methods I-3 / refusal handling: Haiku refuses on 28%/35% of core verb-flip trials and Flash-Lite withholds on 24%. Bias estimates condition on non-refusal after cell exclusion. The paper shows family concentration is not a pure refusal artifact (Flash-Lite ≈0 bias despite high withholding), but does not report whether refusal rates differ systematically by verb, printed order, or label. Differential refusal by frame would make exclusion non-ignorable for the order/lexical split. Please report refusal rates by the crossed factors (or a sensitivity analysis that bounds the bias under plausible missingness) for the high-refusal configurations.","section":"Materials and Methods (I-3, refusal handling); Results Fig. 2"},{"comment":"Results I-2 convergent validity (r = 0.82–0.90 with graded θ): free-choice verdicts are extracted by Claude Opus 4.8, which shares a vendor family with several subjects. The authors flag circularity risk and ship transcripts, but the main-text correlations are presented as primary convergent evidence for the latent scale. Either re-extract a subset with an independent judge (or human coding) and report agreement, or move the I-2 correlations to a clearly caveated secondary check so the θ claim does not rest on same-family judging.","section":"Results “An internal moral scale exists”; Methods I-2"}],"minor_comments":[{"comment":"Fig. 1b: variance-share bars (anchor-direction solid, scale hatched) combine in quadrature; a one-line reminder in the caption that shares are not linearly additive would prevent misreading stacked heights.","section":"Fig. 1b caption"},{"comment":"Notation: b, z, θ, m, s, and σ_repro are introduced across Results and Methods; a short symbol table in Methods or SI would help readers track the [−1, +1] convention and the distinction between descriptive incoherence spreads and bootstrap CIs.","section":"Methods; throughout Results"},{"comment":"The Cheung-verbatim K01 control (order-balanced residual +0.01, CI spanning zero) is important for engaging the source finding; consider elevating one sentence of it into the main Results paragraph on decomposition rather than leaving it mostly in Methods.","section":"Results “The standard readout overlays a format artifact”; Methods “Cheung-verbatim control”"},{"comment":"SI Appendix is heavily referenced for open-weight degeneracy (Nemotron), salt floors, and G coefficients; ensure the main text’s “small open-weight models fail in model-specific ways” is self-contained enough that a reader who skips SI still sees the discrimination-vs-coherence distinction (G vs raw σ_repro).","section":"Results opening; Discussion scopes"},{"comment":"Typos / polish: “theyes–no bias” spacing in the Introduction; occasional missing spaces after em-dashes in the compiled text; confirm that “GPT-5.5” and model snapshot IDs are the intended public names at submission.","section":"Introduction; Methods model panel"}],"recommendation":"minor_revision","confidential_remarks":"The design is unusually careful for this literature; the central identification is sound and the scopes are honest. My minor_revision is driven by abstract–Methods alignment on the logical-bias claim, refusal-by-frame reporting, and the I-2 judge circularity—not by doubt about the order/lexical decomposition. Fit for a methods-forward venue in computational linguistics / AI evaluation is strong. No novelty or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: the big yes-no shift on moral dilemmas is mostly a readout artifact (recency-type order plus pull toward the word \"no\"), not a change in moral stance. Once they cross verb × order × label, the verdict-attached logical channel collapses for every frontier model they test, while graded θ stays nearly format-invariant (σ_repro 0.12–0.21). That is the new result.\n\nWhat they did well is the instrument. Crossed symmetrization is not just more prompts; it turns the confound into named, separable components by construction (Eq. 1). Stance comes from an independent graded battery that never prints yes/no or option order; free-choice I-2 correlates with it at 0.82–0.90; the label-swap identification is clean; matched 12-form baselines kill the obvious corpus artifact; salt replications show the open-weight incoherence is deterministic. The (m, s) logistic is a useful summary, not the load-bearing claim, and they correctly separate s from sampling temperature. Scopes are stated honestly: within-model, these 20 dilemmas, Claude concentration of the large artifact, high-refusal cells excluded. Citation pattern is solid and engages Cheung et al. on shared materials rather than re-litigating their numbers.\n\nSoft spots are real but secondary. The convergent-validity reading of θ is the softest premise; residual anchor-direction floor remains, and absolute s is clip-limited for saturated models. Artifact size is family- and wording-family-dependent, so portability of raw magnitudes is limited. Data/code deposition is promised but DOI not yet live. None of that undercuts the central identification as scoped.\n\nThis is for people who measure model values, safety gates, or survey-style LLM judgments. If you care about whether a binary verdict means what it looks like, read it. It deserves a serious referee; I would accept for peer review and would cite the decomposition and the battery design.","headline":"Clean factorial decomposition of the yes-no bias into order + lexical surface pulls, with logical attachment ~0 under label swap and a coherent graded stance as independent axis.","tokens_in":26769,"tokens_out":523,"would_cite":true,"duration_ms":5225,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The yes-no bias of large language models on moral dilemmas is an artifact of answer order and the word “no,” not a change in moral judgment.","keywords":["large language models","moral judgment","psychometrics","framing effects","AI evaluation","yes-no bias","answer order","crossed symmetrization"],"falsifier":"If the same crossed battery showed frontier models’ graded stance itself swinging as much as the reported binary artifact when only the rating question’s surface form changed, or if fully arbitrary answer labels still produced a large verdict-attached bias, the claim that the scale is format-invariant and the bias is surface-only would fail.","tokens_in":26758,"feed_emoji":"⚖️","tokens_out":933,"duration_ms":15567,"temperature":0.7,"pith_summary":"Frontier language models appear to flip moral verdicts under tiny wording changes, including an amplified yes-no bias that people do not show. This paper argues that the flip is not a change in what the model values. When the same dilemmas are rated on graded scales under many logically equivalent framings, frontier models keep a stable internal stance. The large bias appears only when the answer is forced through yes/no: it splits into a pull toward the last-printed option and a pull toward the word “no.” Replace those words with arbitrary labels and the bias attached to the actual verdict vanishes—the models are not drawn toward rejecting, only toward the printed surface. Measuring what a model values therefore requires crossing the frames of the question, not asking once.","feed_headline":"LLM yes-no bias tracks labels, not moral verdicts","feed_subtitle":"Crossed framings recover a stable stance; the big pull lives in the yes/no readout.","key_machinery":"Crossed symmetrization: every logically irrelevant factor (verb, printed order, answer label, scale, anchor, wording, pole) is flipped in balanced pairs and the factors are crossed so their contributions separate. The symmetric part of each flip-pair estimates stance θ; the antisymmetric part is the artifact. The minimal model P = σ((θ ± m)/s) then summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, distinct from sampling temperature.","core_discovery":"Frontier models carry a coherent internal moral scale: graded ratings of the same dilemmas stay nearly format-invariant under crossed, logically equivalent framings. Forcing the judgment through yes/no overlays a decomposable format artifact—an order bias toward the last-printed option plus a lexical pull toward the word “no”—large mainly in Claude models and smaller under extended reasoning. With arbitrary answer labels the verdict-attached logical bias is approximately zero for every frontier model; the pull follows the printed surface, not the verdict it carries.","pith_inferences":["Safety gates and LLM-as-judge systems that force yes/no may be measuring label and order attachments rather than the intended policy.","The recency-type order bias (opposite classic human primacy) is likely to appear in many multiple-choice readouts, not only moral items.","Interventions that claim to reduce format sensitivity can be audited by tracking m and s rather than raw accuracy alone.","Small models can look “coherent” by being indifferent; any coherence number without a discrimination check can flatter an empty scale."],"forward_implications":["Graded, multi-frame elicitation recovers a model’s moral stance more cleanly than any single forced yes/no.","Single-format binary readouts confound stance with surface format and should not be read as direct measures of value.","Extended reasoning typically shrinks both cross-form incoherence and framing susceptibility where the artifact is large.","The same battery applies unchanged to any dilemma set and binary format, so the decomposition can be rerun on other value domains.","A yes-no bias reported without crossing verb, order, and label cannot be attributed to moral judgment versus surface pull."],"fun_headline_variants":["LLM yes-no bias tracks surface labels and order, not moral verdicts","Frontier models keep stable moral stance under crossed framings","Yes-no artifact is last-option bias plus lexical pull to “no”","With arbitrary labels, verdict bias drops to ~0 across frontier models","Crossed frames show LLMs’ moral scale is nearly format-invariant"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the averaged graded rating under many equivalent framings really is the model’s stable moral stance, so it can serve as the fixed axis against which yes/no artifacts are measured.","fun_headline_variants_meta":{"raw":{"variants":["LLM yes-no bias tracks surface labels and order, not moral verdicts","Frontier models keep stable moral stance under crossed framings","Yes-no artifact is last-option bias plus lexical pull to “no”","With arbitrary labels, verdict bias drops to ~0 across frontier models","Crossed frames show LLMs’ moral scale is nearly format-invariant"]},"model":"grok-4.5","effort":"low","cost_usd":0.00681,"raw_usage":{"total_tokens":1771,"prompt_tokens":954,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":68100000,"prompt_tokens_details":{"text_tokens":954,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":721,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":954,"tokens_out":96,"duration_ms":5635,"temperature":1.0,"reasoning_tokens":721,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T05:53:53.750212+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If the same crossed battery showed frontier models’ graded stance itself swinging as much as the reported binary artifact when only the rating question’s surface form changed, or if fully arbitrary answer labels still produced a large verdict-attached bias, the claim that the scale is format-invariant and the bias is surface-only would fail.","supporting_citations":[],"review_version":1}