Pith. sign in

REVIEW 2 cited by

Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.08514 v2 pith:HVJYBVQG submitted 2025-02-12 cs.CL

classification cs.CL
keywords evaluationfaithfulnessinitialsummaryapproachbeliefdebateerrors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Faithfulness evaluators based on large language models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries. We propose an approach to summary faithfulness evaluation in which multiple LLM-based agents are assigned initial stances (regardless of what their belief might be) and forced to come up with a reason to justify the imposed belief, thus engaging in a multi-round debate to reach an agreement. The uniformly distributed initial assignments result in a greater diversity of stances leading to more meaningful debates and ultimately more errors identified. Furthermore, by analyzing the recent faithfulness evaluation datasets, we observe that naturally, it is not always the case for a summary to be either faithful to the source document or not. We therefore introduce a new dimension, ambiguity, and a detailed taxonomy to identify such special cases. Experiments demonstrate our approach can help identify ambiguities, and have even a stronger performance on non-ambiguous summaries.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    MSumBench is a new English/Chinese benchmark that grades summaries across six domains on faithfulness, completeness, and conciseness, using domain-specific key-facts and multi-agent-debate-assisted human annotations, ...

  2. Beyond Single Models: Enhancing LLM Detection of Ambiguity in Requests through Debate

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A leader-follower multi-agent debate protocol improves ambiguity detection for two of three tested LLMs, but the reported results lack error bars, a clear success metric, and contain internal numerical inconsistencies.

Pith tools