{"id":"0429b133-cad7-4f51-864c-4363206ba054","arxiv_id":"2606.19826","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Honest heterogeneous peers in LLM debates lower harmful revision rates (e.g., 89% to 35%), while adversarial peers raise them (to 90%), and provide defense even against same-family adversaries.","lead":"Experiments show that adding an honest diverse LLM peer reduces harmful answer flips by honest models in debates, while an adversarial peer increases them; this holds even when an adversary is already present. Generalists may read this to learn how diversity can bolster AI system defenses against manipulation.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Attribution of revision-rate shifts to honest vs. adversarial peer type rests on unverified matching of all other panel variables","rationale":"The reader’s weakest_assumption correctly isolates the causal-attribution step as the point where the central claim is least secure. Because the full text is now available, the concrete test above directly checks whether that assumption holds in the reported experiments; until it is performed the UNVERDICTED status remains appropriate.","tokens_in":1846,"tokens_out":382,"duration_ms":21474,"concrete_test":"From the Methods section, extract the exact prompt strings and panel-construction code used for the Llama-3.1-70B MATH-hard runs; compute token counts and message counts for the homogeneous, honest-mixed, and adversarial-mixed conditions. If any differ by >5 % after accounting for model-name substitution, re-execute the homogeneous baseline using the honest-mixed prompt template verbatim and recompute the harmful-revision rate; a shift >10 percentage points would indicate the original comparison is not isolated to peer type.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (89 % → 35 % → 90 % harmful-revision rate for Llama-3.1-70B on MATH-hard) is presented as evidence that heterogeneity itself modulates correction vs. damage. This interpretation requires that the three matched-panel conditions differ only in the identity and honesty label of the added peer; any systematic difference in prompt template, total context length, number of turns, or elicitation phrasing would confound the comparison. The abstract asserts “matched panels” but supplies no quantitative check that token budgets, message counts, or formatting are identical once the peer model name is substituted. If those quantities vary, the observed reversal could be produced by prompt engineering artifacts rather than the peer’s honesty property.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript empirically studies heterogeneous LLM debate by tracking how the presence of an honest or adversarial heterogeneous peer alters honest agents' revision behavior (answer changes that are corrective vs. harmful). It compares matched homogeneous baselines against honest-mixed and adversarial-mixed panels, plus contaminated panels with an existing same-family adversary, across four model families and three reasoning benchmarks. Key reported pattern: an honest peer sharply lowers harmful-revision rates while an adversarial peer reverses the effect; heterogeneity can also reduce loss of initially correct answers when an adversary is already present. Example: for Llama-3.1-70B defenders on MATH-hard, harmful-revision rate falls from 89% (homogeneous) to 35% (honest peer) and returns to 90% (adversarial peer).","tokens_in":1982,"tokens_out":494,"duration_ms":24331,"significance":"If the attribution of rate changes specifically to peer honesty type holds after controls, the work supplies quantitative evidence that heterogeneity functions as both an attack surface and a potential defense in multi-agent LLM systems. The directional consistency across model families and benchmarks, together with the distinction between conditional revision rates and end-of-debate flip rates, offers falsifiable, actionable measurements for protocol design.","major_comments":[{"comment":"Abstract: the headline attribution (89% \to 35% \to 90% harmful-revision rate for Llama-3.1-70B on MATH-hard) requires that the three matched-panel conditions differ only in the identity and honesty label of the added peer. No quantitative verification is supplied that token budgets, message counts, total context length, or formatting remain identical once the peer model is substituted; any systematic difference would confound the comparison and undermine the claim that observed shifts are due to the peer's honesty property rather than prompt artifacts.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract reports specific percentages but does not mention statistical significance tests, confidence intervals, or the number of trials underlying each rate.","section":null},{"comment":"Exact implementation details for the adversarial peer (prompt phrasing, contamination method) and any explicit controls for model-scale or family differences are not summarized at the level needed to assess reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for identifying a methodological point that strengthens the paper. We address the concern below and will revise accordingly.","responses":[{"response":"We agree that explicit quantitative verification is required to isolate the effect of peer honesty. The experimental protocol fixes the debate format, prompt templates, turn structure, and per-turn token caps for all panels; only the model identity in the peer slot changes. However, the manuscript does not report summary statistics (means/variances of tokens per message or total context length) across the three conditions for the headline experiments. We will add a table in the appendix with these statistics for Llama-3.1-70B on MATH-hard (and the other reported settings) to confirm the conditions are matched on these dimensions.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the headline attribution (89% \to 35% \to 90% harmful-revision rate for Llama-3.1-70B on MATH-hard) requires that the three matched-panel conditions differ only in the identity and honesty label of the added peer. No quantitative verification is supplied that token budgets, message counts, total context length, or formatting remain identical once the peer model is substituted; any systematic difference would confound the comparison and undermine the claim that observed shifts are due to the peer's honesty property rather than prompt artifacts."}],"tokens_in":1476,"tokens_out":309,"duration_ms":22713,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is a set of direct measurements on how peer type affects revision behavior in LLM debates. It tracks harmful revision rates and flip rates on initially correct answers across homogeneous, honest-mixed, and adversarial-mixed panels, plus contaminated setups with an existing same-family adversary. The numbers for Llama-3.1-70B on MATH-hard (89% to 35% to 90% harmful revision) and the drop from 31% to 6% flip rate when adding an honest peer are the concrete new data points. The patterns hold directionally across four model families and three benchmarks, which is useful for seeing where magnitude varies with defender strength.\n\nThe work is straightforward empirical tracking of answer changes rather than any derived or fitted claims, so the circularity burden is low. It extends existing debate literature into the adversarial case with specific percentages that were not previously reported.\n\nThe main soft spot is the one flagged in the stress test. The abstract asserts matched panels but gives no quantitative confirmation that token budgets, message counts, or elicitation phrasing stay identical once the peer model is swapped. If those variables shift, the observed reversals could trace to prompt artifacts instead of the honesty label. The abstract also omits any mention of statistical tests, which leaves the consistency claim at a high level only.\n\nThis is the kind of paper that belongs in a reading group focused on multi-agent robustness. It deserves peer review because the measurements are new and falsifiable, even if the controls require closer inspection in revision.","headline":"The paper supplies new quantitative rates showing honest heterogeneous peers cut harmful revisions while adversarial ones restore them, but the matched-panel claim needs explicit verification on prompt and context details.","tokens_in":2474,"tokens_out":390,"would_cite":false,"duration_ms":23405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An honest heterogeneous peer cuts harmful revision rates in LLM debates, while an adversarial peer reverses the reduction.","keywords":["LLM debate","adversarial peers","heterogeneous agents","revision rates","reasoning benchmarks","multi-agent systems","AI safety"],"falsifier":"A controlled replication that keeps every prompt, model scale, and panel size fixed while only swapping the honesty label of the heterogeneous peer and finds no change in harmful-revision rates.","tokens_in":2737,"feed_emoji":"🛡️","tokens_out":700,"duration_ms":20134,"temperature":0.7,"pith_summary":"The paper tests whether diversity among LLM peers in a debate primarily corrects errors or spreads adversarial influence. It does this by comparing how honest agents revise their answers when a new peer is added to homogeneous panels versus panels that already contain an adversary. Across four model families and three benchmarks, the honest peer consistently reduces harmful changes to answers, and the adversarial peer increases them. The same pattern appears when measuring loss of initially correct answers in contaminated panels. This shows that heterogeneity functions as both a potential attack vector and a defense once an adversary is present.","feed_headline":"Honest peer drops LLM harmful revisions from 89% to 35%","feed_subtitle":"Adversarial peers restore the high rate, but the same honest addition also shields correct answers when an adversary is already present.","key_machinery":"Matched panels (homogeneous baseline, honest-mixed, adversarial-mixed) and contaminated panels that track changes in honest agents' revision rates and flip rates when a heterogeneous peer is introduced.","core_discovery":"In matched panels, an honest heterogeneous peer lowers the harmful-revision rate of Llama-3.1-70B defenders on MATH-hard from 89 percent in the homogeneous case to 35 percent, while an adversarial peer returns the rate to 90 percent. The conditional revision rate understates the effect on weak defenders, but the end-of-debate flip rate reveals it. When a same-family adversary is already present, the added honest peer also reduces the rate at which initially correct answers are lost, cutting the flip rate from 31 percent to 6 percent in the same setting. The sign of the effect is stable across families and benchmarks even as its size varies with defender and task difficulty.","pith_inferences":["Debate protocols could deliberately insert diverse honest models as a countermeasure once any adversary is detected.","The same heterogeneity that creates attack surfaces can be turned into a layered defense by adding more than one honest peer type.","Security evaluations of multi-agent LLM systems should routinely include both clean and contaminated panel conditions."],"forward_implications":["Honest heterogeneous peers can protect initially correct answers when a same-family adversary is already present.","The magnitude of the protective effect varies with the defender model and benchmark difficulty.","The pattern of honest-peer benefit and adversarial-peer harm is consistent in sign across model families.","End-of-debate flip rates expose damage that conditional revision rates can hide on weaker defenders."],"fun_headline_variants":["Honest LLM peer cuts harmful revisions from 89% to 35%","Adversarial peer restores harmful revisions to 90%","Honest peer reduces flips on correct answers from 31% to 6%","Heterogeneous peers lower harmful revisions across LLM panels"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That differences in revision rates are caused by the honest or adversarial character of the added peer rather than by prompt wording, model-size differences, or other panel-setup details.","fun_headline_variants_meta":{"raw":{"variants":["Honest LLM peer cuts harmful revisions from 89% to 35%","Adversarial peer restores harmful revisions to 90%","Honest peer reduces flips on correct answers from 31% to 6%","Heterogeneous peers lower harmful revisions across LLM panels"]},"model":"grok-4.3","cost_usd":0.004904,"raw_usage":{"total_tokens":2471,"prompt_tokens":804,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":49037000,"prompt_tokens_details":{"text_tokens":804,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1598,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":804,"tokens_out":69,"duration_ms":12500,"temperature":1.0,"reasoning_tokens":1598,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T17:14:51.753179+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled replication that keeps every prompt, model scale, and panel size fixed while only swapping the honesty label of the heterogeneous peer and finds no change in harmful-revision rates.","supporting_citations":[],"review_version":1}