Pith. sign in

REVIEW 2 cited by

A Proposed S.C.O.R.E. Evaluation Framework for Large Language Models : Safety, Consensus, Objectivity, Reproducibility and Explainability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07666 v1 pith:KUMHYZCN submitted 2024-07-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationframeworkmodelsconsensusexplainabilityhealthcarelanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A comprehensive qualitative evaluation framework for large language models (LLM) in healthcare that expands beyond traditional accuracy and quantitative metrics needed. We propose 5 key aspects for evaluation of LLMs: Safety, Consensus, Objectivity, Reproducibility and Explainability (S.C.O.R.E.). We suggest that S.C.O.R.E. may form the basis for an evaluation framework for future LLM-based models that are safe, reliable, trustworthy, and ethical for healthcare and clinical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems

    cs.MA 2025-05 conditional novelty 6.0 of 10

    A 5,000-prompt medical safety benchmark reveals that decentralized LLM multi-agent teams resist a malicious insider agent better than shared-pool teams, and a personality-screening defense partially restores safety.

  2. Can LLMs faithfully generate their layperson-understandable 'self'?: A Case Study in High-Stakes Domains

    cs.HC 2024-11 reject novelty 5.0 of 10

    The paper defines faithfulness as reproducibility and reports moderate to high reproducibility for LLM-generated layperson algorithms in law, finance, and health, but the algorithms are not shown to reflect the models...

Pith tools