MAJ-EVAL, a document-grounded persona-based multi-agent debate evaluator, correlates more strongly with expert ratings than ROUGE, BERTScore, G-Eval, and ChatEval on children's QA and medical summarization tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
MAJ-EVAL, a document-grounded persona-based multi-agent debate evaluator, correlates more strongly with expert ratings than ROUGE, BERTScore, G-Eval, and ChatEval on children's QA and medical summarization tasks.