REVIEW 1 cited by
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models have advanced clinical Natural Language Generation, creating opportunities to manage the volume of medical text. However, the high-stakes nature of medicine requires reliable evaluation, which remains a challenge. In this narrative review, we assess the current evaluation state for clinical summarization tasks and propose future directions to address the resource constraints of expert human evaluation.
Forward citations
Cited by 1 Pith paper
-
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
MAJ-EVAL, a document-grounded persona-based multi-agent debate evaluator, correlates more strongly with expert ratings than ROUGE, BERTScore, G-Eval, and ChatEval on children's QA and medical summarization tasks.
Discussion (0). Sign in to comment.