REVIEW 2 cited by
A Proposed S.C.O.R.E. Evaluation Framework for Large Language Models : Safety, Consensus, Objectivity, Reproducibility and Explainability
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A comprehensive qualitative evaluation framework for large language models (LLM) in healthcare that expands beyond traditional accuracy and quantitative metrics needed. We propose 5 key aspects for evaluation of LLMs: Safety, Consensus, Objectivity, Reproducibility and Explainability (S.C.O.R.E.). We suggest that S.C.O.R.E. may form the basis for an evaluation framework for future LLM-based models that are safe, reliable, trustworthy, and ethical for healthcare and clinical applications.
Forward citations
Cited by 2 Pith papers
-
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
A 5,000-prompt medical safety benchmark reveals that decentralized LLM multi-agent teams resist a malicious insider agent better than shared-pool teams, and a personality-screening defense partially restores safety.
-
Can LLMs faithfully generate their layperson-understandable 'self'?: A Case Study in High-Stakes Domains
The paper defines faithfulness as reproducibility and reports moderate to high reproducibility for LLM-generated layperson algorithms in law, finance, and health, but the algorithms are not shown to reflect the models...
Discussion (0). Continue with ORCID to comment.