{"id":"71bbda2c-28f6-478f-96df-b0831c7dd41c","arxiv_id":"2505.21112","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Two AI ethics debates with differently composed panels reached the same policy recommendation but through different arguments and coalitions, showing that panel membership can shift reasoning even when the facts are fixed.","lead":"This paper introduces ADEPT, a system that stages structured debates between AI personas representing different ethical perspectives to deliberate on medical dilemmas. It tests whether changing which moral viewpoints are present changes how an AI panel reasons about allocating scarce ventilators.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that panel composition materially changes deliberation is undercut by lack of replication: with one temperature-0.7 run per panel and no seeds or controls, observed vote and theme shifts cannot be separated from stochastic LLM variation.","rationale":"I agree with the reader that the single-run, no-replication design is the load-bearing weakness. The paper is transparently framed as a proof-of-concept, the code and data are public, and the qualitative analysis is human-verified, which are genuine strengths. However, the central empirical claim—that changing the moral perspectives on the panel materially changed deliberation—requires that the two observed debates differ because of the persona substitution and not because of sampling noise. That requirement is not met by one temperature-0.7 run per panel. The authors explicitly list 'only two debate iterations' as a limitation, so the concern is internal to the paper, not a purely external standard. The reader's proposed remedy (replications, toned-down abstract, statistical treatment of stochastic variation) is exactly what the evidence needs. The final policy outcome being unchanged further narrows the claim to argumentation and coalitions, which are more volatile across runs. Therefore the conditional verdict remains appropriate: the contribution is a promising methodology, but the evidence for composition-driven effects requires replications and a more careful statement of what changed. No verdict change is warranted beyond the reader's CONDITIONAL, so I mark UNCHANGED.","tokens_in":21509,"tokens_out":3927,"duration_ms":50407,"concrete_test":"Run each of the two panel configurations at least 10 times under identical prompts, varying only the random seed (or sampling temperature draws), and code each transcript for final votes and for presence/absence of target themes (moral injury, legal risk, public trust) using a blinded protocol. Then compute a permutation test or bootstrap confidence interval on the difference in vote distributions and theme prevalence between panel conditions. If within-condition variation across runs is comparable to or larger than the between-condition difference, the observed shifts in Debate 2 cannot be attributed to panel composition. Also report whether the number of continuing personas whose final option changed is stable across runs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's causal inference—that replacing two personas redirected the debate and changed continuing personas' positions—rests on a single run per panel (§3.2.1). The engine uses temperature=0.7 (§3.1), so every generation step is stochastic; the debate is a multi-turn chain in which early random differences can propagate into later votes and justifications. With N=1 per condition, the observed differences in coalitions and themes are statistically indistinguishable from run-to-run noise. This is not a mere call for more data: the abstract claims that moral perspectives 'can materially change the outcome,' yet the final policy outcome was identical in both debates (Option 2, 4–2); only supporting coalitions and justifications changed, which are precisely the aspects most sensitive to sampling variation. The paper's own §5.3 acknowledges 'only two debate iterations' and the possibility that base-model inclinations drove the consistent majority, lending support to this concern. Additionally, the abstract states that 'four continuing personas' changed final positions, but the vote tables show the Disability-Rights Advocate stayed with Option 2 across both debates; only three continuing personas changed their option. This discrepancy further weakens the reported effect and suggests the qualitative framing may overstate the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ADEPT, an LLM-orchestrated multi-agent framework that stages structured ethical debates among personas defined by explicit ethical-theory and stakeholder specifications. It demonstrates the system on a ventilator-triage scenario using two six-persona panels that differ only in two members: Debate 1 includes a Catholic Bioethicist and a Care Ethicist, while Debate 2 substitutes a Deontologist and a Legal Arbiter. Both debates return the same majority policy outcome (Option 2, Clinical + Equity Weighted Lottery, by 4–2), but the supporting coalitions, argumentative themes, and some individual votes differ. The paper claims three contributions: a transparent, replicable workflow; evidence that panel composition can materially change outcomes even when factual inputs are fixed; and an analysis of implications and future directions for AI-mediated ethical deliberation.","tokens_in":21651,"tokens_out":3052,"duration_ms":35791,"significance":"If the causal claim about panel composition were adequately supported, ADEPT would be a useful and genuinely transparent tool for exploring normative pluralism, with clear value for ethics pedagogy, policy prototyping, and the study of deliberative dynamics. The paper's strengths include publicly available code and full debate transcripts, detailed persona specifications in Appendix B, direct quotation from the generated transcripts, and an unusually candid limitations section that identifies epistemic-reliability and black-box concerns. However, the central empirical contribution is not established by the current evidence: the comparison rests on a single stochastic run per panel condition, and several headline statements in the abstract and conclusion overstate what the data show. The paper is best read as a qualitative proof of concept, and the revision should either add the repeated-run evidence needed for the causal claim or explicitly scale the claim back to that scope.","major_comments":[{"comment":"The paper's central causal claim—that replacing two personas redirected the debate and changed continuing personas' positions—rests on one debate run per panel at temperature 0.7, with no seeds, repeated runs, or statistical controls. Because every generation step is stochastic and the debate is a multi-turn chain, the observed differences in coalitions and thematic emphasis are indistinguishable from run-to-run noise. This concern is load-bearing for contribution (ii), which asserts that moral perspectives 'can materially change the outcome.' As written, the manuscript does not provide sufficient evidence for that claim; it would need multiple runs per condition with reported variability, or a clear reframing of the result as a single illustrative case study rather than evidence of a systematic effect.","section":"§3.1, §3.2.1, Tables 3–4"},{"comment":"The abstract states that the altered membership 'changed four continuing personas' final positions' and that the work provides 'evidence that the moral perspectives included in such panels can materially change the outcome.' This is not what the reported data show: the final policy outcome is identical in both debates (Option 2, 4–2), and the vote tables indicate that the Disability-Rights Advocate voted for Option 2 in both debates, so only three continuing personas changed their option. The abstract and Section 4.1 should be corrected to state the actual outcome and the actual number of vote changers, and the language of 'changed the outcome' should be replaced with a more precise description of changed coalitions and justifications.","section":"Abstract, §4.1, Tables 3–4"},{"comment":"The limitations section acknowledges that the study includes 'only two debate iterations' and that the consistent majority for Option 2 'could be influenced by subtle, uninstructed inclinations within the foundational model itself.' This directly undercuts the conclusion's claim that the comparative analysis revealed 'tangible shifts in ... the final policy preferences of the simulated committee.' The conclusion should be aligned with the limitations: at most, the paper shows that a single pair of runs produced different argumentative trajectories, not that panel composition reliably shifts deliberative outcomes. This is a logical inconsistency between the paper's explicit caveats and its summary claims.","section":"§5.3, §7"}],"minor_comments":[{"comment":"The phrase 'opening statements → rebuftals → secret ballot' contains a typo; 'rebuftals' should be 'rebuttals.'","section":"§3.1"},{"comment":"The subsection numbering jumps from 3.2.2 to 3.2.4, with no 3.2.3; the numbering should be corrected or a placeholder section added.","section":"§3.2"},{"comment":"The table titles contain the typo 'Vote Talley'; this should read 'Vote Tally.'","section":"Tables 3 and 4"},{"comment":"The phrase 'how varying ethical perspectives influences debate dynamics' should use the plural verb 'influence' to agree with 'perspectives.'","section":"§3.2.4"},{"comment":"The illustrative examples are drawn only from Debate 1, which limits their usefulness for the comparative claim; including at least one parallel exchange from Debate 2 would let readers assess the claimed shift in argumentative style directly.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is an honest proof-of-concept with a valuable open-source artifact, but the headline empirical claim is not supported by the current single-run design. I recommend a revision that either adds repeated runs with reported variance or explicitly reframes the paper as a qualitative case study. If the author chooses the latter, the abstract and conclusion must be revised accordingly, since several statements currently overstate the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuine proof-of-concept for using LLM personas in structured bioethics deliberation, and it is unusually honest about its own limits. The code, YAML persona definitions, and full debate transcripts are public. That transparency is real and should count for something. The qualitative analysis of the two debates is also careful, with human verification of quotes and explicit acknowledgment of confabulated citations. As a demonstration that multi-agent debate can be repurposed for ethical deliberation with auditable outputs, it does what it claims.\n\nWhat is new is narrow but real: swapping a Catholic bioethicist and care ethicist for a deontologist and legal arbiter changed the argumentative themes and the voting coalition, even though the final policy option stayed the same. The paper frames this as evidence that panel composition 'can materially change the outcome,' which overstates it. The policy outcome did not change. The themes and justifications changed, which is interesting, but the abstract's phrasing invites a stronger reading than the data support.\n\nThe load-bearing weakness is the one the stress-test note identifies, and I think it lands. There is one run per panel at temperature 0.7, no seeds, no repetition. With a multi-turn stochastic chain, the observed differences in coalition and framing are statistically indistinguishable from run-to-run noise. The paper's own Section 5.3 concedes only two debate iterations. A single replication per condition would not settle it, but a few runs per panel with temperature sweeps would at least show the effect is not a coin flip. I also checked the vote tables: the Disability-Rights Advocate voted Option 2 in both debates, so the abstract's claim that 'four continuing personas' changed final positions is inaccurate; it is three. That is a minor wording issue, but it matters because the whole paper hinges on getting the shift right.\n\nThe author anticipates the epistemic concerns better than most: confabulation, lack of fact-checking, black-box reasoning, and the risk that the base model's own inclinations drove the consistent majority are all discussed. That does not rescue the central inference, but it makes the paper a fair and usable starting point rather than an overclaiming one. The citation pattern is fine, and the self-citations point to public artifacts. There is no sign of fitted parameters or unacknowledged tuning; this is not a fitting-as-prediction problem, it is an interpretation problem.\n\nWho is this for? People working on AI-mediated deliberation, bioethics pedagogy, or policy prototyping will get value from the workflow and the candid limitations. It deserves a serious referee, and I would accept it for review despite my skepticism about the central claim. My recommendation: engage with it, but treat the panel-composition effect as a hypothesis to test, not a demonstrated result.","headline":"A transparent, honest proof-of-concept for LLM persona deliberation panels, but the central causal claim about panel composition rests on a single stochastic run per condition.","tokens_in":22241,"tokens_out":666,"would_cite":true,"duration_ms":13985,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that ADEPT, a system of LLM personas debating a ventilator-triage dilemma, shows that changing which ethical perspectives sit on the panel materially changes how the debate unfolds and how votes land, even when the facts…","keywords":["LLM personas","multi-agent debate","bioethics deliberation","ventilator triage","normative pluralism","auditable workflow","ethical simulation"],"falsifier":"Run both panels repeatedly, say twenty times each with different random seeds, and compare how often each continuing persona changes its vote. If the four shared personas shift their positions just as often when the panel membership is unchanged as when it is swapped, the claim that panel composition drives the outcome would be falsified.","tokens_in":21219,"feed_emoji":"⚖️","tokens_out":7509,"duration_ms":76106,"temperature":0.7,"pith_summary":"ADEPT is a proof-of-concept workflow in which large language model personas, each embodying a distinct ethical framework or stakeholder role, debate a fixed policy question in three phases and then vote. The paper argues that this workflow is a transparent, replicable way to study moral deliberation, and it presents as its main evidence a controlled comparison: two six-person panels debated the same ventilator-allocation scenario with identical facts and options, differing only in that a Catholic Bioethicist and a Care Ethicist were replaced by a Deontologist and a Legal Arbiter. Both panels chose the same majority policy, a clinically weighted lottery that avoids withdrawing ventilators for reallocation, but the second panel reached it through different arguments and a different coalition, with four continuing personas changing their final positions. The author reads this as evidence that the moral perspectives included in such a panel can materially change deliberative outcomes, and positions ADEPT as a tool for exploring normative pluralism, ethics education, and policy prototyping.","feed_headline":"Two AI ethics panels reached the same policy by different coalitions","feed_subtitle":"Swapping a Catholic bioethicist and care ethicist for a deontologist and legal arbiter shifted four votes.","key_machinery":"The load-bearing mechanism is ADEPT's three-phase deliberation loop: opening statements, rebuttals, and a secret ballot, each logged for audit. Personas are defined by a structured YAML schema that fixes their ethical principle, approach, core questions, decision criteria, deliberation style, forbidden moves, and citations, so each panel member acts as a 'moral lens' rather than a generic debater. The controlled input is a fixed scenario with four pre-specified allocation options, and the experimental instrument is the two-panel comparison: four personas are shared, two are swapped, and every prompt, response, and vote is recorded as an inspectable artefact.","core_discovery":"On its own terms, the paper's discovery is that panel composition is a decisive variable in simulated ethical deliberation. With every factual input held fixed, swapping two of the six personas redirected the debate's attention toward moral injury, legal risk, and public trust, and it changed the final positions of the four personas who appeared in both debates: the Front-Line ICU Nurse and the Consequentialist moved from Option 1 to Option 2, while the Virtue Ethicist moved from Option 2 to Option 3. The stated upshot is that ADEPT offers 'a transparent, replicable workflow for running and analysing multi-agent AI debates in bioethics' and evidence that 'the moral perspectives included in such panels can materially change the outcome even when the factual inputs remain constant.'","pith_inferences":["Because each panel ran once at temperature 0.7, the observed differences could in principle be random model variation; running the same panels many times with different seeds would test whether the vote shifts are reliably caused by composition.","A natural extension is to use ADEPT-style panels as an 'ethical red team' for draft policies, comparing the objections AI personas raise with those of real ethics committees before a guideline is adopted.","Comparing the same personas across different underlying language models would separate persona-level effects from base-model leanings, which the paper flags as an open question."],"forward_implications":["If ADEPT works as claimed, deliberative bodies could probe how the same clinical facts yield different ethical recommendations depending on which moral perspectives are represented.","The auditable transcript and vote log would let ethicists, regulators, and clinicians trace exactly which arguments carried which votes, rather than only seeing a final recommendation.","Replacing personas can surface concerns that would otherwise stay implicit, such as legal exposure under human-rights law or the risk that prognosis scores disadvantage disabled patients.","A single majority policy can hide divergent justifications, so consensus in AI deliberation should be reported together with the coalition that produced it."],"supporting_citations":[{"why":"Provides the ADEPT code, configuration files, and debate outputs that make the workflow and its transcripts publicly inspectable.","marker":"[21]"},{"why":"Records Debate 1's full transcript and votes, serving as one half of the controlled comparison.","marker":"[22]"},{"why":"Records Debate 2's full transcript and votes, serving as the other half of the controlled comparison.","marker":"[23]"},{"why":"Supplies the fair-allocation ethical framework that informs the ventilator-triage scenario and options.","marker":"[24]"},{"why":"Provides a state-level framework for scarce mechanical ventilation allocation that the four options are modeled on.","marker":"[25]"},{"why":"Grounds the scenario in UK crisis-standards guidance.","marker":"[26]"},{"why":"Supplies the multi-agent conversation architecture that ADEPT's orchestrator adapts.","marker":"[3]"},{"why":"Provides the multi-agent debate approach and its critique, motivating ADEPT's transparent logging and fixed-option design.","marker":"[4]"},{"why":"Demonstrates LLM-mediated deliberation on policy questions, the line of work ADEPT extends to structured ethical panels.","marker":"[5]"}],"fun_headline_variants":["AI ethics panels reach same verdict despite swapped personas","Swapping two AI personas changes four votes, not the policy","LLM debate panels: same outcome, different moral paths","Panel makeup shifts AI ethics debate, not final choice","Two AI ethics panels, one policy, four changed minds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the differences between the two debates were caused by swapping the two personas, rather than by random variation in the language model, since each panel was run only once at temperature 0.7.","fun_headline_variants_meta":{"raw":{"variants":["AI ethics panels reach same verdict despite swapped personas","Swapping two AI personas changes four votes, not the policy","LLM debate panels: same outcome, different moral paths","Panel makeup shifts AI ethics debate, not final choice","Two AI ethics panels, one policy, four changed minds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000376,"raw_usage":{"total_tokens":2028,"prompt_tokens":991,"completion_tokens":1037,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":958}},"tokens_in":607,"tokens_out":1037,"duration_ms":8065,"temperature":1.0,"reasoning_tokens":958,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:34:53.579966+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run both panels repeatedly, say twenty times each with different random seeds, and compare how often each continuing persona changes its vote. If the four shared personas shift their positions just as often when the panel membership is unchanged as when it is swapped, the claim that panel composition drives the outcome would be falsified.","supporting_citations":[{"cited_title":"ADEPT [Internet]","cited_arxiv_id":null,"evidence_quote":"Provides the ADEPT code, configuration files, and debate outputs that make the workflow and its transcripts publicly inspectable."},{"cited_title":"ADEPT Debate 1 [Internet]","cited_arxiv_id":null,"evidence_quote":"Records Debate 1's full transcript and votes, serving as one half of the controlled comparison."},{"cited_title":"ADEPT Debate 2 [Internet]","cited_arxiv_id":null,"evidence_quote":"Records Debate 2's full transcript and votes, serving as the other half of the controlled comparison."},{"cited_title":"Fair Allocation of Scarce Medical Resources in the Time of Covid-19","cited_arxiv_id":null,"evidence_quote":"Supplies the fair-allocation ethical framework that informs the ventilator-triage scenario and options."},{"cited_title":"Too Many Patients…A Framework to Guide Statewide Allocation of Scarce Mechanical Ventilation During Disasters","cited_arxiv_id":null,"evidence_quote":"Provides a state-level framework for scarce mechanical ventilation allocation that the four options are modeled on."},{"cited_title":"Community health services two-hour urgent community response standard [Internet]","cited_arxiv_id":null,"evidence_quote":"Grounds the scenario in UK crisis-standards guidance."},{"cited_title":"AI can help humans find common ground in democratic deliberation","cited_arxiv_id":null,"evidence_quote":"Demonstrates LLM-mediated deliberation on policy questions, the line of work ADEPT extends to structured ethical panels."}],"review_version":1}