On a small set of NTCIR-18 reports, GPT-based evaluators, especially GPT-Black, tracked expert judgments of causal medical explanations better than similarity metrics such as BERTScore and cosine similarity.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
On a small set of NTCIR-18 reports, GPT-based evaluators, especially GPT-Black, tracked expert judgments of causal medical explanations better than similarity metrics such as BERTScore and cosine similarity.