{"id":"b5ce618d-e2b6-4d82-aade-ef17f1348e99","arxiv_id":"2510.10002","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"In multi-agent debates over everyday moral dilemmas, GPT-4.1 almost never revises in simultaneous settings but conforms strongly in sequential settings, while Claude 3.7 and Gemini 2.0 Flash revise far more often.","lead":"This paper runs multi-agent debates between three large language models on 1,000 moral dilemmas and finds that deliberation format changes how willing models are to revise their verdicts and which values they invoke. GPT-4.1 rarely changes its mind in simultaneous debate but strongly conforms in sequential debate, showing interaction protocol shapes moral judgment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reader's single-run concern lands on the derived layers, not raw CoV rates: at N=1000 the 0.6% vs 41.2% gap has tight binomial CIs, but Table 1's inertia/conformity decomposition and the LLM-judge value statistics are single-sample and unvalidated.","rationale":"The reader flagged single-run stochasticity as the weakest assumption. In good faith I checked whether the headline numbers could move: with N=1,000 per condition, the observed CoV rates and one-round consensus rates are binomial proportions with ±1–3% 95% CIs (e.g., GPT sync 0.6% ≈ [0.2,1.3]%, Gemini sync 41.2% ≈ [38,44]%; round-robin one-round 90% vs 40%, each ±~2%). A re-run at temperature 1 cannot plausibly flip the qualitative orderings, so the reader's specific mechanism is overstated. The single-run design is nonetheless load-bearing where the paper goes beyond model-free rates: (i) the pooled multinomial model (Eq. 2) used for the headline 'inertia vs conformity' synthesis (Table 1) is fit with ~4,000 dilemma fixed effects and unspecified 'weak L2 regularization', without convergence or model-checking diagnostics; its CIs assume between-round independence within a deliberation, which temperature-1 sampling violates (verdict flips are per-dilemma-correlated), and for GPT-second round-robin data, γ_within,GPT and α_GPT can absorb the same conform-then-stick behavior, so the 'inertial in sync / most conformist in round-robin' decomposition is not tested for identifiability; (ii) the value statistics (Figs 4–7) rest on single-shot Gemini 2.5 Flash classifications with no inter-rater reliability, so 'inherited values' and the value-similarity/consensus link have unquantified judge noise. Independent support: the order-reversal design and the model-free consensus rates already carry the central qualitative claim, and the authors' Section 5 limitation statement is transparent. I therefore keep the reader's CONDITIONAL verdict (UNCHANGED), but the condition should specifically require repeated-run/robustness validation of the Eq. 2 parameters and the value-classification layer, not the raw CoV rates.","tokens_in":25066,"tokens_out":18409,"duration_ms":173660,"concrete_test":"Repeated-sampling protocol: take a random 200-dilemma subset, re-run all conditions (3 synchronous pairs, 6 head-to-head round-robin orders, 6 three-way orders) 5 times at temperature 1 (5 seeds), and refit Eq. 2 on each repetition. Compute cross-repetition 95% intervals for α_GPT, γ_prev,GPT, γ_within,GPT and the Claude/Gemini analogues. If any repetition's γ_within,GPT overlaps Claude's γ_within, or the GPT-vs-Gemini ordering of γ_within flips, then Table 1's conformity synthesis is not stable and the single-run design is a genuine threat to the quantitative claims; if all parameters stay within ~20% of Table 1 across repetitions, the concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central qualitative claim is likely robust: with N=1,000 per condition the aggregate rates are binomial with tight CIs — GPT sync CoV 0.6% (95% CI ≈ [0.2,1.3]%) vs Gemini 41.2% ([38,44]%), round-robin one-round consensus 90% vs 40% by order (±2%). Re-running at temperature 1 will not flip these orderings, so the reader's stated mechanism — CoV rates shifting substantially — is quantitatively implausible. The single-run design is load-bearing exactly where the paper leaves model-free rates behind. First, Eq. 2 pools all protocols into one multinomial model with ~4,000 dilemma fixed effects φ_dv and unspecified weak-L2 regularization; no convergence, identifiability, or model-check diagnostics are reported. Table 1's CIs treat each round as independent, but verdict flips at temperature 1 are serially correlated within a deliberation (flip propensity is per-dilemma), so the reported precision is overstated. In round-robin data where GPT goes second, its first verdict is already conditioned on the opponent's verdict, so γ_within,GPT and α_GPT can absorb the same 'conform-then-stick' behavior; nothing in the pooled fit demonstrates that the inertial-in-sync / most-conformist-in-round-robin split is a stable decomposition rather than a pooling artifact. Second, the value layer is likewise one-shot: Gemini 2.5 Flash classifications (§3.3) are single draws with no inter-rater reliability, so Figs 4–7 and the 'inherited values drive verdict changes' claim carry unquantified judge noise. The paper's synthesis, Table 1, is the least secure link; Section 5's hedge that 'aggregate results are likely robust' covers the raw rates but not the model fits or value statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares two multi-agent deliberation protocols—synchronous and round-robin—for three proprietary LLMs (GPT-4.1, Claude 3.7 Sonnet, Gemini 2.0 Flash) across 1,000 everyday moral dilemmas from r/AmItheAsshole. It reports change-of-verdict (CoV) rates, consensus rates, value-usage and value-inheritance patterns, and fits a multinomial logistic model to separate 'inertia' and 'conformity' parameters. The central finding is that deliberation format changes LLMs' willingness to revise moral verdicts and that this is model-dependent: GPT shows high inertia in synchronous settings but high conformity in round-robin settings, while Claude and Gemini behave differently. The paper also tests a modified system prompt aimed at steering consensus-seeking behavior.","tokens_in":25598,"tokens_out":3072,"duration_ms":31689,"significance":"If the findings hold, they establish that interaction protocol is a first-order variable in multi-agent moral judgment rather than a secondary implementation detail, with implications for deployed multi-agent systems. The study uses a large, naturalistic dilemma corpus, releases code, and reports bootstrapped confidence intervals. The aggregate CoV differences (e.g., 0.6% vs. 41.2%) are large and internally consistent. However, the manuscript's central qualitative claim is more robust than its derived quantitative decomposition: the raw CoV rates have tight binomial confidence intervals despite single-run data, whereas the fitted inertia/conformity parameters and the value-level analyses rest on single-sample, unvalidated measurements.","major_comments":[{"comment":"The single-run design is acknowledged in §5 ('we ran each experiment once'). This is not fatal for the headline CoV rates: for N=1,000, the binomial 95% CI for 0.6% is roughly [0.1%, 1.1%] and for 41.2% roughly [38.2%, 44.2%], so the aggregate ordering is robust. However, the multinomial model in Eq. (2) estimates α and γ parameters from those same single runs and treats each round as independent. Verdict flips at temperature 1 are serially correlated within a deliberation, and the ~4,000 dilemma fixed effects φ_dv are not accompanied by convergence, identifiability, or model-check diagnostics. Consequently, Table 1's odds ratios (α_GPT 8.27, γ_within,GPT 8.68) may overstate precision and the inertia/conformity split could absorb the same 'conform-then-stick' behavior. Please provide repeated runs (or a variance decomposition), cluster-robust standard errors, and posterior predictive or","section":"§5, §4.1, Table 1, Eq. (2)"},{"comment":"The value classification uses a single LLM judge (Gemini 2.5 Flash) with a single pass per response, no inter-rater reliability, and no repeated sampling. This is load-bearing for the claims that certain values are 'inherited' and that value alignment drives consensus. Since the judge is itself a Gemini-family model and the deliberators include Gemini 2.0 Flash, systematic judge bias could affect the value-occurrence and value-similarity results. The reported bootstrapped CIs only capture variation over dilemmas, not variation over judge draws or judge identity. Please add multi-judge or repeated-judge measurements, agreement statistics (e.g., Cohen's κ or a judge-replication analysis), and show that the value-level conclusions are stable under judge variation.","section":"§3.3, Figs. 4–7"},{"comment":"The comparison of CoV rates between synchronous and round-robin settings is not apples-to-apples. In round-robin, the second mover's Round 1 verdict is already conditioned on the opponent's verdict, so the model's 'conformity' (γ_within) and its 'inertia' (α) can both reflect the same behavior of accepting the prior verdict and then sticking to it. The pooled model in Eq. (2) does not by itself demonstrate that the inertial-in-synchronous / conformist-in-round-robin split is a stable decomposition rather than an artifact of pooling protocols with different information structures. Please report first-round agreement rates separately from later-round revision rates, and ideally fit the model separately for each protocol or include protocol-specific interactions with explicit exposure timing.","section":"§4.3, Fig. 2c–d, Table 1"}],"minor_comments":[{"comment":"The notation nprev_vd and nwithin_vd,r is not fully defined: what counts as a 'previous round' when multiple models are present, and how are counts normalized across rounds of different lengths? Clarify the summation indices and the exact coding for round-robin vs. synchronous settings.","section":"Eq. (2)"},{"comment":"The caption says 'b. The fraction of deliberations where a specific value was inherited,' but the panel structure labels a–c as 'Difference in Value Occurrence' and d–f as inherited values. The text in §4.2 also refers to panels in an order that does not match the caption. Please reconcile.","section":"Fig. 4 caption"},{"comment":"The value list contains 'Environmental consciousness' twice. Deduplicate.","section":"Appendix F"},{"comment":"No model fit statistics (e.g., log-likelihood, AIC, or a null-model comparison) are reported for Eq. (2). Reporting these would help readers judge whether the inertia/conformity decomposition actually improves fit over a model with only baseline and fixed effects.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's honest limitation statement is a strength, but it also identifies the exact gaps that need to be closed. The headline qualitative finding (protocol shapes moral judgment, with GPT behaving differently in synchronous vs. round-robin) is likely robust given the large N and tight aggregate CIs. The main risk is over-interpretation of the fitted parameters and the value-judge results. I would not reject: the fixes—repeated runs or cluster-robust inference, model diagnostics, judge-reliability analysis, and a clearer separation of first-round vs. later-round effects—are within the scope of a revision and do not require a fundamentally new experimental paradigm."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper has a robust central finding—deliberation format changes how GPT-4.1, Claude 3.7, and Gemini revise moral verdicts—and a shakier set of fitted parameters trying to explain it. I read it as a useful empirical contribution, not a sealed one.\n\nWhat's new: most multi-agent debate work uses verifiable tasks or single-turn AITA judgments. This paper runs 1,000 AITA dilemmas in two protocols, synchronous and round-robin, and shows a qualitative reversal: GPT is almost immovable in synchronous settings (0.6–3.1% CoV) but becomes the most conforming model in round-robin (within-round odds ratio 8.68). The consensus and order-effects results are also cleanly presented. The paper is transparent about methods, supplies bootstrapped CIs, includes system prompts in the appendix, and flags its own limitations. The steering experiment showing that a \"balanced\" prompt moves GPT's CoV by 5–18x but doesn't restore consensus is a nice touch.\n\nSoft spots. The single-run design is the least of it. With 1,000 dilemmas, raw CoV rates have tight binomial CIs; re-running at temperature 1 may shift things but won't flip 0.6% vs 41.2%. The real problem is the derived layer. Equation 2 fits one multinomial model with roughly 4,000 dilemma fixed effects, weak L2 regularization, and no convergence or identifiability diagnostics. The reported CIs treat rounds as independent, but verdict flips within a deliberation are serially correlated per dilemma. More importantly, in round-robin where GPT goes second, its first-round verdict is already conditioned on the opponent's verdict, so gamma_within and alpha can both absorb the same \"conform then stick\" behavior. The sync-inertia / round-robin-conformity split might be a pooling artifact rather than a stable decomposition. The value layer has the same one-shot problem: Gemini 2.5 Flash classifications are single draws with no inter-rater reliability, so Figures 4–7 and the \"inherited values\" claims carry unquantified judge noise. The paper's own Section 5 hedge about robustness covers aggregate rates, not Table 1 or the value statistics.\n\nWho should read it: people designing multi-agent systems for sensitive advice, and anyone working on sociotechnical alignment. It deserves a serious referee—the design is thoughtful, the limitations are stated, and the central claim is strong enough to survive replication. A referee should ask for repeated sampling of at least a subset of dilemmas, identifiability checks on the fitted model, and a small human or second-judge validation of the value classification. I'd take it to reading group.","headline":"The headline effect is real and worth taking seriously, but the inertia/conformity decomposition is less secure than it looks.","tokens_in":26064,"tokens_out":2562,"would_cite":true,"duration_ms":25179,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that whether language models revise their moral verdicts in multi-agent debate depends less on which model they are than on how the debate is orchestrated: parallel ('synchronous') exchanges make GPT-4.1 nearly immovable, w","keywords":["multi-agent debate","moral judgment","verdict revision","conformity","inertia","interaction protocol","value alignment","LLM deliberation"],"falsifier":"Re-run the 1,000-dilemma experiment, say 10 times per dilemma at temperature 1, and compute change-of-verdict rates and the fitted α and γ parameters with bootstrapped confidence intervals. If GPT-4.1's synchronous change-of-verdict rate overlaps Claude's, or if its round-robin within-round conformity odds ratio falls below Gemini's, the paper's central claim that protocol flips model flexibility would be falsified.","tokens_in":24979,"feed_emoji":"⚖️","tokens_out":4505,"duration_ms":40184,"temperature":0.7,"pith_summary":"This paper claims that how large language models are made to talk to each other—in parallel or in sequence—systematically changes which models change their moral verdicts, and that this behavior is not a fixed model trait. Across 1,000 everyday Reddit dilemmas, it finds GPT-4.1 is almost immovable in synchronous debate, changing verdicts in just 0.6–3.1% of cases, but becomes the most conformist model in round-robin debate, with a within-round conformity odds ratio of 8.68. Claude and Gemini show the opposite pattern: they are highly flexible in parallel settings but less conformist when exposed to prior verdicts sequentially. If correct, the paper establishes interaction protocol as a first-order variable in multi-agent moral judgment, as important as which model is used. This matters for anyone building agentic systems that give moral advice, because the same model can appear stubborn or sycophantic depending on how the conversation is structured.","feed_headline":"Debate format flips which AI changes its moral verdict","feed_subtitle":"GPT-4.1 rarely revises in parallel debate but conforms hard in sequential rounds.","key_machinery":"The central mechanism is a multinomial logistic model: for each model m, dilemma d, and round r, the log-odds of a verdict depend on that model's baseline preferences, a fixed dilemma effect, an inertia term α_m for repeating its own previous verdict, and two conformity terms γ_prev,m (frequency of a verdict in earlier rounds) and γ_within,m (frequency within the current round, non-zero only in round-robin). The two deliberation protocols—synchronous (parallel, simultaneous responses) and round-robin (sequential, with later models seeing earlier verdicts)—are the manipulated variable that makes inertia and conformity separable. Values are classified via a curated 48-value taxonomy and measur","core_discovery":"The authors claim that verdict revision and consensus in LLM deliberation are governed by two opposing forces—inertia (repeating one's own prior verdict) and conformity (yielding to verdicts seen from others)—and that the balance of these forces is format-dependent. In synchronous debate, GPT-4.1's change-of-verdict rate is 0.6–3.1% while Claude and Gemini revise 28–41% of the time; in round-robin debate, GPT-4.1 and Gemini conform strongly to the verdict they see first, with GPT-4.1's within-round conformity odds ratio at 8.68. Consensus is more likely when models' invoked values converge, and value similarity (Jaccard index over a 48-value taxonomy) rises by 30–60% in deliberations that re","pith_inferences":["An implication the paper leaves implicit: a moral-advice system that uses synchronous parallel debates may overstate a model's confidence, while a round-robin system may silently inherit the first model's bias—so a hybrid protocol could serve as a calibrating middle ground.","Because the paper runs each dilemma once at temperature 1, the headline odds ratios are point estimates; a natural following experiment would repeat the same 1,000 dilemmas across many seeds to measure variance and confirm the model-ordering results.","The value-inheritance result (Claude and Gemini often inherit GPT's personal-autonomy values, while GPT rarely inherits empathy values) hints that influence flows asymmetrically—worth testing whether that asymmetry persists if models exchange system prompts or are prompted to match the other's style.","An untested extension: head-to-head and three-way formats produced different consensus rates; larger panels (five or more agents) or mixed human-AI panels could reveal whether the format effect saturates or reverses."],"forward_implications":["Consensus in multi-agent moral deliberation often reflects conformity or inertia rather than genuine persuasion, so 'consensus' is not evidence of correctness.","The same model can flip from rigid to compliant merely by switching the interaction protocol, so system designers should test values under both formats before deployment.","Value convergence closely tracks verdict convergence: when models agree they share roughly three of five values, and consensus-reaching deliberations raise Jaccard similarity by 30–60%.","System prompt steering ('balance consensus and correctness') increases verdict changes but does not raise consensus rates, meaning models often move to different verdicts rather than converging.","Order effects in round-robin deliberation are large enough to shift the final verdict distribution—for example, GPT steering over 70% of dilemmas to NTA when going first in a three-way debate."],"fun_headline_variants":["Debate style decides which AI budges on morals","Parallel debate makes AI stubborn, sequential makes them conform","Moral verdicts flip based on how AI debaters speak","Study: Debate format shifts which AI revises verdicts","Round-robin debate bends AI verdicts; parallel debate locks them"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Every headline number rests on a single run per dilemma at a temperature of 1; because language-model output is stochastic, a repeat run could shift the revision rates and the fitted inertia/conformity parameters, possibly changing the relative ordering of models.","fun_headline_variants_meta":{"raw":{"variants":["Debate style decides which AI budges on morals","Parallel debate makes AI stubborn, sequential makes them conform","Moral verdicts flip based on how AI debaters speak","Study: Debate format shifts which AI revises verdicts","Round-robin debate bends AI verdicts; parallel debate locks them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001022,"raw_usage":{"total_tokens":4223,"prompt_tokens":896,"completion_tokens":3327,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":3245}},"tokens_in":640,"tokens_out":3327,"duration_ms":21116,"temperature":1.0,"reasoning_tokens":3245,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T10:20:17.757958+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 1,000-dilemma experiment, say 10 times per dilemma at temperature 1, and compute change-of-verdict rates and the fitted α and γ parameters with bootstrapped confidence intervals. If GPT-4.1's synchronous change-of-verdict rate overlaps Claude's, or if its round-robin within-round conformity odds ratio falls below Gemini's, the paper's central claim that protocol flips model flexibility would be falsified.","supporting_citations":[],"review_version":1}