{"id":"ef7f35dc-6911-405a-96f5-115a7de70268","arxiv_id":"2507.08440","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.","lead":"This paper builds a multi-agent system in which large language models debate a policy issue and a judge agent decides when the participants have reached agreement. The authors test six LLMs on stance benchmarks and on a simulated drug-policy conference, and argue that adding an agreement detector makes the simulated debate cover the same ground as a real expert conference.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim depends on ChatGPT4-as-judge scoring without task-specific human validation; the same model is also a top contestant, so reliable agreement detection in debates is not yet established.","rationale":"In good faith, the paper makes a plausible contribution: it evaluates six LLMs on two stance benchmarks, shows that open and smaller models are competitive, and proposes a modular multi-agent architecture for simulating decision conferences with a judge agent. The objective stance-detection results are informative, and the qualitative transcript comparison is a reasonable illustration of the system's behavior. However, the central claim is specifically about agreement detection in dynamic, multi-turn debates, and the only direct evaluation of that capability is ChatGPT4-as-a-judge scoring of a small set of extracted decision points. The absence of task-specific human validation is not a mere cosmetic issue: without it, the paper cannot distinguish genuine agreement-detection ability from the evaluator's priors, formatting preferences, or self-preference. The with/without-judge comparison is also a single unblinded transcript, so the claimed improvement in topic coverage is not robustly established. This is a missing-evidence problem rather than an internal contradiction, and it is addressable with a focused annotation study. Because the reader already conditioned acceptance on exactly this kind of validation, my read does not change the verdict.","tokens_in":20144,"tokens_out":4140,"duration_ms":52998,"concrete_test":"Have at least two annotators with decision-conference expertise independently label every judge-agent decision point from the simulated transcripts (at minimum the 30 scored in Table 5) as either 'agreement reached' or 'debate should continue,' using the criteria of the real drug-policy conference. Compute inter-annotator agreement and compare the majority human labels with ChatGPT4's scores and with the resulting model ranking; for example, report Cohen's kappa between ChatGPT4 labels and human labels. If kappa is below 0.6, or if ChatGPT4's ranking of the six models does not match the human-based ranking, the central claim that LLMs reliably detect agreement in dynamic debates is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims that 'LLMs can reliably detect agreement even in dynamic and nuanced debates.' For this to hold, the judge agent's agreement decisions in the simulated conferences must be correct. The direct evidence for correctness is Table 5: ChatGPT4, acting as an LLM-as-a-judge, grades five extracted decision points per model on a 1-10 scale. Two problems arise. First, the evaluator is itself an LLM with no human annotations for this specific task; the cited validation (reference [11]) concerns general chat-assistant evaluation on MT-Bench, not agreement detection in multi-agent decision conferences, and that reference also documents judge biases. Second, ChatGPT4 is itself one of the six judged models and receives a perfect 10/10 in Table 5, creating a self-preference risk. The objective benchmarks in Section 5.1 measure stance detection and stance polarity on isolated texts, not the relational decision of whether two debating agents have reached agreement over a dialogue, so they cannot validate the judge's core function. The with/without-judge comparison in Section 5.2.2 rests on one manually inspected transcript, and the claimed difference (coverage of the 'public' cluster) could be stochastic. Thus the central reliability claim currently rests on an unvalidated and partially self-referential evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-agent system for simulating decision conferences, built on AutoGen, in which a moderator guides participant agents through staged debates and a judge agent decides whether the participants have reached agreement. The judge agent is then evaluated in two ways: objectively, by testing six LLMs (Gemma 2 9B, Gemma 7B, Mixtral 8x7B, LLaMA 3 70B, ChatGPT 3.5 Turbo, ChatGPT 4) on stance detection and stance polarity detection using the VAST and Claim Stance Classification datasets; and subjectively, by using ChatGPT 4 as an LLM-as-a-judge to score the judge agent's decisions in a simulated drug-policy decision conference. The authors also compare a simulation with the judge agent against one without it, arguing that the judge agent leads to more complete topic coverage (specifically, covering the 'public' cluster). The paper concludes that LLMs can reliably detect agreement in dynamic debates and that incorporating an agreement-detection agent improves the quality and realism of simulated deliberations.","tokens_in":20320,"tokens_out":2776,"duration_ms":35187,"significance":"If the central claim holds, the paper would provide a practical blueprint for using zero-shot LLMs as agreement detectors in multi-agent deliberation, with a potentially useful application in expert elicitation and decision-support systems. The objective evaluation is a genuine strength: it compares six LLMs against established fine-tuned baselines on standard benchmarks, demonstrating that mid-sized open models such as Gemma 2 9B and LLaMA 3 70B can be competitive without task-specific training. The architectural description and the reproducible benchmark comparisons are also valuable. However, the central claim about 'reliable agreement detection in dynamic and nuanced debates' rests primarily on a subjective evaluation whose ground truth is itself an LLM judgment (ChatGPT 4), and the paper does not provide task-specific human validation for that judgment. The with/without-judge comparison is based on a single transcript. These gaps currently make the main conclusion stronger than the evidence supports.","major_comments":[{"comment":"The subjective evaluation uses ChatGPT 4 as the LLM-as-a-judge to grade each model's judge-agent decisions, but no human annotations are provided for agreement detection in decision conferences. The paper justifies this choice by citing reference [11] (MT-Bench/Chatbot Arena), yet that work validates LLM judges on general chat-assistant quality, not on detecting agreement in multi-agent debates, and it explicitly documents judge biases such as self-preference and verbosity bias. The fact that ChatGPT 4 receives a perfect 10/10 in its own row of Table 5 is therefore not evidence of correctness, and the scores cannot currently support the abstract's claim that LLMs 'reliably detect agreement.' A task-specific validation set, or at least a random sample of judge decisions scored by human annotators, is needed before the central claim can be accepted.","section":"Section 5.2.1, Table 5"},{"comment":"In the VAST evaluation, the neutral class contains only 2 examples, and after reporting that all models essentially fail on this class (F1 scores of 0.0 to 0.028), the paper states that this 'can be considered less impactful' and that the primary focus should remain on pro and con. This is a post-hoc dismissal of a class that is directly relevant to the system's purpose: the judge agent must distinguish 'agreement,' 'disagreement,' and 'still debating' (a neutral state). Excluding or downweighting the neutral class changes the reported macro-F1 substantially, and the paper should either justify the exclusion a priori or report micro-averaged metrics and per-class results without the post-hoc reinterpretation.","section":"Section 5.1.1, Table 2"},{"comment":"The stance polarity results show a systematic bias toward predicting negative polarity across all models, with the paper noting that 'some positive labels are predicted as negative.' The rationalization that this 'is not that bad' because missing negative polarity would cause premature termination is not supported by the task definition: if the judge agent relies on polarity to detect agreement, then misclassifying positive (supportive) statements as negative could cause the debate to continue unnecessarily or cause an agreement to be missed. The paper should present a confusion matrix or error analysis for the polarity task and discuss how this bias affects the judge agent's agreement decisions, rather than asserting that the failure mode is benign.","section":"Section 5.1.3, Table 4 and Figure 3"},{"comment":"The comparison of the system with and without the judge agent is based on a single manually inspected transcript for one topic. The claimed benefit is that the with-judge simulation covers all seven thematic clusters (health, social, political, public, crime, economic, cost), while the without-judge simulation misses 'public.' Because LLM simulations are stochastic (no random seed reporting, no repeated runs, no error bars), a single run cannot establish that this difference is due to the judge agent rather than to sampling variability. The paper should report multiple runs (e.g., 5–10 per condition) with a measure of coverage variability, or explicitly frame the result as an illustrative case study rather than evidence for the general claim that the judge agent 'prevents premature transitions between topics.'","section":"Section 5.2.2"},{"comment":"The objective evaluation operationalizes agreement detection as stance detection and stance polarity detection on isolated benchmark texts. However, the judge agent's actual function in the simulated decision conference is a relational, dialogue-level decision: given the exchange between two or more participants, determine whether they have reached agreement. Stance classification of individual claims is a necessary component but not sufficient evidence for the ability to perform this relational judgment, because agreement detection requires tracking whether a later utterance aligns with, responds to, and resolves prior statements. The paper should either provide a dialogue-level objective evaluation (e.g., on a conversational agreement or negotiation dataset) or explicitly narrow the central claim to stance-based agreement detection rather than 'agreement detection in dynamic and nuanced debates.'","section":"Section 4.1 and Section 5.1"}],"minor_comments":[{"comment":"Several typographical and formatting issues should be corrected: 'V AST' appears with an inconsistent space (e.g., 'VAST' vs 'V AST'), '1.355' should be '1,355' in Section 4.1.1, and 'T able 1' in Section 5.1 has an erroneous space.","section":"Throughout"},{"comment":"The pseudocode labels an 'evaluation agent' that scores the debate, while the text mostly refers to a 'judge agent' for agreement detection. The relationship between these two agents should be clarified, since the evaluation agent appears to use LLM-as-a-judge during the simulation while the judge agent makes the agreement decision.","section":"Section 3.2.1, Algorithm 1"},{"comment":"The paper states that all models perform exceptionally well on the Claim Stance Classification dataset and that the top three surpass the 0.849 state-of-the-art baseline, but the comparison could be made more precise by reporting the variance or confidence intervals, especially since zero-shot prompting can be sensitive to prompt wording.","section":"Section 5.1.2"},{"comment":"The paper states 'Not applicable' for code availability, but the custom speaker selection function and system prompts are central to the reproducibility of the simulations. Making at least the prompts and the speaker selection logic available would strengthen the paper.","section":"Section 7.1.7 'Code availability'"},{"comment":"The use of 'five decisions from the judge agent' is not justified in the text; the paper should explain why five decision points were selected, how they were sampled, and whether the judge agent produced more than five decisions that were excluded.","section":"Section 5.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is interesting and the objective benchmark results are a useful contribution, but the central claim about reliable agreement detection is currently supported mainly by an LLM-as-a-judge evaluation that lacks task-specific human validation and by a single-transcript comparison. The authors should be encouraged to add a small human-annotated evaluation of judge decisions or a validated dialogue-level benchmark, and to run the with/without-judge comparison multiple times. This is fixable within the scope of the manuscript, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper applies established pieces — zero-shot stance detection, AutoGen multi-agent debate, and LLM-as-a-judge — to a new target: simulating decision conferences and using a dedicated judge agent to detect agreement. That combination is genuinely new, and the objective benchmark section is the strongest part. Six models are evaluated on VAST and the Claim Stance dataset, and the results show that zero-shot LLMs, including smaller open models, beat several trained baselines. The single case study is also useful: adding the judge agent made the simulated debate cover the 'public' cluster that the without-judge run missed.\n\nThe soft spots are real but addressable. The central claim in the abstract — LLMs 'reliably detect agreement in dynamic and nuanced debates' — is stronger than the evidence. The subjective evaluation uses ChatGPT4 as the judge of whether the judge agent detected agreement, and ChatGPT4 is also one of the six judged models, scoring 10/10 throughout. The cited validation (MT-Bench) covers general chat evaluation, not decision-conference agreement detection, so there is a self-preference risk and a task-transfer gap. The with/without-judge comparison is one transcript; one missed cluster could be stochastic. The objective benchmarks measure stance detection on isolated texts, not the relational judgment of whether two debating agents have reached agreement over a dialogue, so they do not directly validate the judge's core function. The paper also handles the VAST neutral class (2 examples) and the systematic negative-polarity bias via post-hoc rationalization rather than squarely addressing the metric problems. No code or prompts are released, which hampers reproduction.\n\nTo be fair, the paper is more careful in the conclusion than in the abstract: it explicitly says this must be explored with more decision conferences and notes the judge does not evaluate accuracy or relevance. But the abstract still oversells.\n\nWho gets value: researchers working on multi-agent deliberation simulation, LLM-as-judge methodology, and decision support. They will find the architecture description and the benchmark comparison useful, though they should not treat the reliability claim as settled.\n\nMy verdict: this deserves a serious referee, but with the expectation of major revision. I would ask for human or at least independent validation of a sample of judge decisions, multiple simulation runs, and either code/prompt release or error bars. If those are added, the paper could be a solid contribution.\n\nRecommendation: send to peer review, conditional on revision.","headline":"Useful new application with solid objective benchmarks, but the central reliability claim is currently carried by a self-referential LLM-as-a-judge setup and a single transcript.","tokens_in":20881,"tokens_out":2433,"would_cite":false,"duration_ms":29512,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Zero-shot LLMs can reliably detect agreement in multi-agent debates, and inserting a dedicated agreement-detection judge into a simulated decision conference makes the simulated debate cover the same ground as a real expert conference.","keywords":["Large Language Models","Agreement Detection","Multi-Agent Collaboration","Decision Conference","Stance Detection","LLM-as-a-Judge","Group Decision-Making"],"falsifier":"Have human experts rate the judge agent's agreement calls on the simulated transcripts and compare their ratings with ChatGPT 4's; if the two diverge on this task, the central claim loses its measurement ground. Simpler still, run the with- and without-judge simulations on a second real decision conference with a published outcome: if the without-judge debate also covers all the real conference's topic clusters, the single transcript comparison that supports the judge agent's benefit would no longer distinguish the two designs.","tokens_in":19898,"feed_emoji":"🤝","tokens_out":9488,"duration_ms":91376,"temperature":0.7,"pith_summary":"This paper builds a multi-agent system that simulates a decision conference, a structured meeting where experts debate a complex issue and work toward consensus, and asks whether large language models can detect when the debating participants have reached agreement. The authors evaluate six LLMs zero-shot on stance detection and stance polarity detection benchmarks, then place the models in a 'judge agent' role that decides whether the agents should keep debating or move on. They report that the best models, ChatGPT 4, LLaMA 3 70B, and the smaller open-source Gemma 2 9B, detect agreement reliably even in nuanced, multi-round debate, matching or beating task-specific models trained for stance detection. The central evidence for the system is a comparison against a real drug-policy decision conference: without the judge agent, the simulated debate covered six of the seven thematic clusters of criteria from the real conference; with the judge agent, it covered all seven, matching the real outcome. If correct, this means zero-shot LLMs can serve as agreement detectors in deliberation systems, and a dedicated agreement-detection module makes simulated group decision-making behave more like the real thing.","feed_headline":"With an agreement-detecting agent, simulated debates hit all 7 clusters","feed_subtitle":"An LLM judge that spots premature agreement lets simulated debates cover the same ground as real expert panels.","key_machinery":"The load-bearing component is the judge agent: an LLM that, after each round of participant debate, decides whether agreement has been reached and signals the moderator, through a custom speaker-selection function, to either continue debating or advance to the next stage of the conference. The judge's task is operationalized by two NLP benchmarks: stance detection, which identifies whether a statement supports or opposes a proposition, and stance polarity detection, which classifies sentiment as positive, negative, or neutral; these serve as the objective proxy for agreement. The simulated conference itself is built on the AutoGen framework for multi-agent conversation, and its output is evaluated by a second LLM-as-a-judge layer plus a manual transcript comparison against a published real-world decision conference on drug policy. The architecture keeps the judge's verdict binary, agreement or continued debate, so the system's value rests entirely on whether that binary call is made at the right moment.","core_discovery":"The central claim is that LLMs can perform zero-shot agreement detection in dynamic, nuanced debates, and that a dedicated agreement-detection agent materially improves a simulated decision conference. On objective benchmarks, the top LLMs match or surpass task-specific stance-detection systems without any fine-tuning or prompt engineering; the three leaders from the benchmarks, LLaMA 3 70B, Gemma 2 9B, and ChatGPT 4, remain the top performers when placed inside the simulated conference and judged by an independent LLM-as-a-judge evaluation. The authors' most concrete evidence for the judge agent's value is a direct outcome comparison: a simulated debate about drug-policy criteria, run without the judge, produced six of the seven thematic clusters (health, social, political, public, crime, economic, cost) identified by real experts, omitting 'public'; the same debate with the judge detected that agreement was premature and continued until all seven clusters were covered, reproducing the real conference's outcome. The paper concludes that agreement detection is a critical component for LLM-based simulation of group decision-making, and that open-source models of moderate size are sufficient for the role.","pith_inferences":["The objective evidence measures stance, not agreement itself; the step from 'a statement supports or opposes a claim' to 'two debating agents have reached agreement' is an assumption the benchmarks never directly test, so a dedicated agreement-annotation dataset would be the natural next experiment.","The with-judge benefit rests on a single transcript; a statistical test over many simulated conferences, with varied topics, personas, and participant counts, is needed to confirm that the judge agent, rather than prompt randomness, produces the fuller coverage.","Because the judge only checks that agreement has been reached, not that the agreed content is accurate or grounded, retrieval-augmented grounding of participant claims could change both how fast agreement forms and whether it is well-founded.","The paper's result that a 9-billion-parameter model matches GPT-4 suggests agreement detection may hinge on instruction-following and output-format compliance more than raw reasoning scale, a prediction testable by varying prompt strictness."],"forward_implications":["Zero-shot LLMs can replace fine-tuned stance-detection models for agreement detection in debate settings, removing the need for task-specific training data.","Adding a judge agent that detects agreement prevents premature transitions between debate topics, so simulated discussions cover the full range of perspectives a real expert panel would raise.","Mid-sized open-source models perform at the level of ChatGPT 4 for this task, so agreement-detection systems can be run locally and at lower cost.","Debate-based multi-agent systems generally can use a dedicated agreement-detection module to know when to stop arguing and consolidate a decision, improving both efficiency and coverage.","The simulation approach could support real decision-making workflows, such as expert elicitation workshops, by revealing which perspectives a group is at risk of overlooking."],"supporting_citations":[{"why":"Supplies the V AST stance-detection benchmark and the trained classical baselines (CMaj, BoWV, C-FFNN, BiCond, Bert-joint, TGA-NET) that the zero-shot LLMs are measured against.","marker":"[37]"},{"why":"Supplies the Claim Stance Classification Dataset, the 0.849 state-of-the-art accuracy baseline, and the stance-polarity labels used for the second objective task.","marker":"[38]"},{"why":"Justifies the LLM-as-a-judge evaluation method; the paper relies on it for the claim that ChatGPT 4 has the highest agreement with human evaluators, which grounds the subjective evaluation.","marker":"[11]"},{"why":"The published drug-policy decision conference whose outcome (27 criteria in seven thematic clusters) serves as the reference standard for the with/without-judge comparison.","marker":"[41]"},{"why":"The AutoGen framework in which the simulated decision conference is implemented, providing the agent infrastructure and conversation management.","marker":"[33]"},{"why":"The theoretical account of decision-conference structure and facilitation that the simulation's stages, moderator role, and judge role are designed to mirror.","marker":"[29]"}],"fun_headline_variants":["LLM judge agent nails agreement detection in simulated debates","Agreement-detecting LLM makes simulated debates match real panels","Zero-shot LLM detects agreement, improves debate simulation","Agreement agent boosts simulated conferences to full coverage","LLMs spot premature agreement in multi-agent decision talks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes ChatGPT 4's scores are a trustworthy measure of whether the judge agent correctly detected agreement, but the evidence for that trustworthiness comes from general chatbot-quality assessment, not from agreement detection in decision conferences.","fun_headline_variants_meta":{"raw":{"variants":["LLM judge agent nails agreement detection in simulated debates","Agreement-detecting LLM makes simulated debates match real panels","Zero-shot LLM detects agreement, improves debate simulation","Agreement agent boosts simulated conferences to full coverage","LLMs spot premature agreement in multi-agent decision talks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000602,"raw_usage":{"total_tokens":2848,"prompt_tokens":1020,"completion_tokens":1828,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":1761}},"tokens_in":636,"tokens_out":1828,"duration_ms":12444,"temperature":1.0,"reasoning_tokens":1761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:18:55.725279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human experts rate the judge agent's agreement calls on the simulated transcripts and compare their ratings with ChatGPT 4's; if the two diverge on this task, the central claim loses its measurement ground. Simpler still, run the with- and without-judge simulations on a second real decision conference with a published outcome: if the without-judge debate also covers all the real conference's topic clusters, the single transcript comparison that supports the judge agent's benefit would no longer distinguish the two designs.","supporting_citations":[{"cited_title":"In: Lapata, M., Blunsom, P., Koller, A","cited_arxiv_id":null,"evidence_quote":"Supplies the Claim Stance Classification Dataset, the 0.849 state-of-the-art accuracy baseline, and the stance-polarity labels used for the second objective task."},{"cited_title":"International Journal of Drug Policy 56, 144–152 (2018)","cited_arxiv_id":null,"evidence_quote":"The published drug-policy decision conference whose outcome (27 criteria in seven thematic clusters) serves as the reference standard for the with/without-judge comparison."},{"cited_title":"Advances in Decision Analysis, 375–399 (2007) https://doi.org/10.1017/cbo9780511611308.020","cited_arxiv_id":null,"evidence_quote":"The theoretical account of decision-conference structure and facilitation that the simulation's stages, moderator role, and judge role are designed to mirror."}],"review_version":1}