MAJ-EVAL, a document-grounded persona-based multi-agent debate evaluator, correlates more strongly with expert ratings than ROUGE, BERTScore, G-Eval, and ChatEval on children's QA and medical summarization tasks.
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
Large Language Models have advanced clinical Natural Language Generation, creating opportunities to manage the volume of medical text. However, the high-stakes nature of medicine requires reliable evaluation, which remains a challenge. In this narrative review, we assess the current evaluation state for clinical summarization tasks and propose future directions to address the resource constraints of expert human evaluation.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
MAJ-EVAL, a document-grounded persona-based multi-agent debate evaluator, correlates more strongly with expert ratings than ROUGE, BERTScore, G-Eval, and ChatEval on children's QA and medical summarization tasks.