Pith. sign in

REVIEW 5 cited by

Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11595 v3 pith:LEASKPAM submitted 2023-05-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsdebatecollaborationinter-consistencyvariouscollaborateconsensuseffectively
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown impressive capabilities in various applications, but they still face various inconsistency issues. Existing works primarily focus on the inconsistency issues within a single LLM, while we complementarily explore the inter-consistency among multiple LLMs for collaboration. To examine whether LLMs can collaborate effectively to achieve a consensus for a shared goal, we focus on commonsense reasoning, and introduce a formal debate framework (FORD) to conduct a three-stage debate among LLMs with real-world scenarios alignment: fair debate, mismatched debate, and roundtable debate. Through extensive experiments on various datasets, LLMs can effectively collaborate to reach a consensus despite noticeable inter-inconsistencies, but imbalances in their abilities can lead to domination by superior LLMs. Leveraging a more advanced LLM like GPT-4 as an authoritative judge can boost collaboration performance. Our work contributes to understanding the inter-consistency among LLMs and lays the foundation for developing future collaboration methods. Codes and data are available at https://github.com/Waste-Wood/FORD

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

    cs.MA 2025-07 conditional novelty 5.0 of 10

    A leader LLM trained with a GRPO variant that conditions on frozen agent responses improves both collaborative and zero-shot accuracy on BBH, MATH, and MMLU.

  2. Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.

  3. Agentic Feature Augmentation: Unifying Selection and Generation with Teaming, Planning, and Memories

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A router-selector-generator LLM agent team with offline PPO and dual memories unifies feature selection and generation, reporting improved downstream performance on six tabular datasets.

  4. DRF: LLM-AGENT Dynamic Reputation Filtering Framework

    cs.AI 2025-09 conditional novelty 4.0 of 10

    DRF combines an LLM rating network, a reputation update rule, and a UCB-style selection strategy to filter low-quality LLM agents during multi-agent task execution, reporting improved pass@1 and lower simulated cost o...

  5. AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need

    cs.CL 2025-06 reject novelty 4.0 of 10

    A divide-and-conquer multi-agent framework with task forests and specialized roles improves math and code benchmarks but not commonsense or domain QA, and the adaptive heterogeneous-LLM engine is never tested.

Pith tools