REVIEW 4 cited by
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong correlation with human judges in the context of general instruction following. However, for instructions that require specialized knowledge, the validity of using LLMs as judges remains uncertain. In our study, we applied a mixed-methods approach, conducting pairwise comparisons in which both subject matter experts (SMEs) and LLMs evaluated outputs from domain-specific tasks. We focused on two distinct fields: dietetics, with registered dietitian experts, and mental health, with clinical psychologist experts. Our results showed that SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health when evaluating overall preference. Additionally, the results indicated variations in SME-LLM agreement across domain-specific aspect questions. Our findings emphasize the importance of keeping human experts in the evaluation process, as LLMs alone may not provide the depth of understanding required for complex, knowledge specific tasks. We also explore the implications of LLM evaluations across different domains and discuss how these insights can inform the design of evaluation workflows that ensure better alignment between human experts and LLMs in interactive systems.
Forward citations
Cited by 4 Pith papers
-
Benchmarking LLMs on File System Design and Implementation
LLMs pass 88–96% of basic file-system tasks but only 38–42% of optimization and new-feature tasks on the new 505-task Phi-Bench benchmark.
-
InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation
An interactive tree visualization of LLM sampling lets evaluators cover the same harmful-response space as random sampling with up to 5x fewer samples.
-
Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs
A Gita-based mental-health dialogue dataset helps small LLMs score higher on spirituality-oriented metrics, but the evaluation loop is largely self-referential.
-
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
MedReadCtrl instruction-tunes LLaMA3 to control readability at 12 grade levels, reporting lower readability errors than GPT-4 and higher content scores on unseen clinical simplification.
Discussion (0). Sign in to comment.