REVIEW 34 cited by
Debating with More Persuasive LLMs Leads to More Truthful Answers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information. The method we evaluate is debate, where two LLM experts each argue for a different answer, and a non-expert selects the answer. We find that debate consistently helps both non-expert models and humans answer questions, achieving 76% and 88% accuracy respectively (naive baselines obtain 48% and 60%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.
Forward citations
Cited by 34 Pith papers
-
Does Multi-Agent Debate Improve AI Feedback on Research Papers?
Authors of economics meta-analyses found a single-pass AI report more useful than two multi-agent debate tools that cost up to thirty times more to run.
-
More Is Not More: What Matters for Diversity in LLM Opinions?
Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.
-
Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages
FMG uses GPT-4o's image and text understanding to make chemical decisions inside a clique-tree algorithm, producing interpretable molecular grammars that generate valid, class-specific molecules from tiny datasets.
-
$\Sigma$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
Online symmetric reliability memory for LLM multi-agent systems accumulates bounded competence and peer-relationship evidence and supports steering, routing, and weighted voting without retraining.
-
It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation
In a 98-participant fact-checking study, the rhetorical style of AI advice changed accuracy, confidence, and preference, with step-by-step explanations helping most and user preference diverging from performance.
-
DeliChess: A Multi-party Dialogue Dataset for Deliberation in Chess Puzzle Solving
Groups that deliberate on chess puzzles improve their move quality on average, but more probing questions only increase outcome variance unless they focus on concrete answer options.
-
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis
Engineered personas plus IS/OOS evidence retrieval make cheap multi-model panels produce tested claim maps and expose RLHF-induced blind spots, including asymmetric AI-risk challenge.
-
Human-AI Complementarity: A Goal for Amplified Oversight
Confidence-based routing of fact-verification to humans, plus evidence-only AI assistance, beats either human or AI raters alone: 91.3% hybrid accuracy vs 87.7% for the AI rater.
-
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
A structured multi-agent LLM debate grounded in simulation is proposed for co-designing robot morphology and reward, but the body reports only a single Ant experiment.
-
Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
MCSU-based vocabulary alignment plus distance-based dynamic selection (DDS) lets several LLMs vote token-by-token, beating single models and prior ensemble baselines on multiple reasoning benchmarks without training.
-
When to Trust Context: Self-Reflective Debates for Context Reliability
SR-DCR uses an asymmetric debate plus self-confidence to gate whether a model follows context or its prior, improving ClashEval accuracy on several models.
-
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
Teach2Eval ranks LLMs by the average accuracy improvement of weak student models they teach, and claims this ranking correlates more strongly with Chatbot Arena and LiveBench than direct evaluation does.
-
An alignment safety case sketch based on debate
The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices fo...
-
Debate Helps Weak-to-Strong Generalization
Debate transcripts from two strong models, used as context when training an ensemble of weak models, improve weak-to-strong generalization on four NLP classification benchmarks.
-
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
DPS uses partial-image descriptions as supervision to prompt a vision-language model to correct itself, reducing attack success by about 76% across six datasets.
-
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
A benchmark of six long-horizon game environments shows current LLMs and VLMs struggle on hard tasks and often do worse when given images.
-
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
ForestBench evaluates LLM multi-agent systems by mapping execution traces to collaboration graphs and measuring their match against per-query reference forests of verified-success structures.
-
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
Judge upgrades are not interchangeable: only Qwen3 1.7B→4B yields robust adjacent gains, MiniMax adjacent releases do not, and stronger judges reduce but do not remove bias or correlated jury errors.
-
How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs
A leader LLM trained with a GRPO variant that conditions on frozen agent responses improves both collaborative and zero-shot accuracy on BBH, MATH, and MMLU.
-
Collaboration among Multiple Large Language Models for Medical Question Answering
An iterative collaboration framework where three LLMs exchange summarized reasoning on disagreed questions raises USMLE-style answer accuracy by 5.2 to 6.6 percentage points per model and raises consensus from 51% to 83%.
-
AUTOLAW: Enhancing Legal Compliance in Large Language Models via Case Law Generation and Jury-Inspired Deliberation
AutoLaw's verifier-ranked legal-role jury with a similar-case demonstration beats majority voting for violation detection on three law and policy benchmarks.
-
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
A multi-agent, rubric-based meta-judge pipeline improves the precision of selecting correct LLM judgments on JudgeBench from 61.71% to 77.26%, with majority voting.
-
Engineering AI Judge Systems
A constitution-based, search-driven development framework for AI judge systems improves judged accuracy by up to 6.2% on commit message generation, with about 58% of general principles reused across five languages.
-
DS@GT ARC at Touch\'e: Large Language Models for Retrieval-Augmented Debate
Frontier LLM judges show strong within-family agreement in a retrieval-augmented debate task, but that consensus does not reliably predict official human-annotation F1, with Quality showing the largest gap.
-
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
AI has rapidly crossed expert baselines on bounded tasks on a still-jagged frontier, so human work must shift from production to specification, verification, and oversight.
-
Hierarchical Debate-Based Large Language Model (LLM) for Complex Task Planning of 6G Network Management
A hierarchical debate framework, in which LLMs first decompose a 6G task and then refine each sub-task, improves keyword coverage over one-shot and regular single-level debate on the 6GPlan benchmark.
-
Information Bargaining: Bilateral Commitment in Bayesian Persuasion
Bayesian persuasion is restated as a two-sided bargaining game, but the proof reduces to a relabeling and the empirical validation is circular.
-
Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings
DEEVO evolves better LLM prompts by debating outputs and selecting survivors with Elo ratings, without requiring labeled data or a hand-written fitness function.
-
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.
-
Language Games as the Pathway to Artificial Superhuman Intelligence
A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.
-
DS@GT at Touch\'e: Large Language Models for Retrieval-Augmented Debate
In the Touché 2025 retrieval-augmented debate task, LLM debaters generated verbose but relevant responses, and LLM evaluators were strict and only moderately consistent.
-
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
Test-time scaling techniques such as long chain-of-thought, tree search, and self-correction define the paper's 'cognition engineering' paradigm, which it surveys, taxonomizes, and tutorials.
-
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
A survey of scalable oversight for superalignment, reviewing weak-to-strong generalization, debate, RLAIF, and sandwiching, and concluding that current methods are not yet sufficient for superintelligent AI.
-
Algebraic Evaluation Theorems
Under an error-independence assumption, the decisions of three binary classifiers determine their true accuracies up to exactly two alternative solutions, and one of them is the true evaluation.
Discussion (0). Continue with ORCID to comment.