Pith. sign in

REVIEW 2 cited by

Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04472 v2 pith:TGNF4JDN submitted 2024-08-08 cs.CL

classification cs.CL
keywords debateagent4debatecompetitiveframeworkhumanhumansllmsagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Competitive debate is a complex task of computational argumentation. Large Language Models (LLMs) suffer from hallucinations and lack competitiveness in this field. To address these challenges, we introduce Agent for Debate (Agent4Debate), a dynamic multi-agent framework based on LLMs designed to enhance their capabilities in competitive debate. Drawing inspiration from human behavior in debate preparation and execution, Agent4Debate employs a collaborative architecture where four specialized agents, involving Searcher, Analyzer, Writer, and Reviewer, dynamically interact and cooperate. These agents work throughout the debate process, covering multiple stages from initial research and argument formulation to rebuttal and summary. To comprehensively evaluate framework performance, we construct the Competitive Debate Arena, comprising 66 carefully selected Chinese debate motions. We recruit ten experienced human debaters and collect records of 200 debates involving Agent4Debate, baseline models, and humans. The evaluation employs the Debatrix automatic scoring system and professional human reviewers based on the established Debatrix-Elo and Human-Elo ranking. Experimental results indicate that the state-of-the-art Agent4Debate exhibits capabilities comparable to those of humans. Furthermore, ablation studies demonstrate the effectiveness of each component in the agent structure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simulating Ethics: Using LLM Debate Panels to Model Deliberation on Medical Dilemmas

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Two AI ethics debates with differently composed panels reached the same policy recommendation but through different arguments and coalitions, showing that panel membership can shift reasoning even when the facts are fixed.

  2. Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A structured multi-agent debate framework with domain-specialized AI agents and a five-dimension scoring rubric improves LLM-based fake news detection by several F1 points.

Pith tools