A new German open-response medical benchmark shows LLM judges can match physician agreement (κ≈0.69 vs 0.71) yet lack clinical caution and show lineage-dependent scoring bias.
Natural Language Processing of Referral Letters for Machine Learning–Based Triaging of Patients With Low Back Pain to the Most Appropriate Intervention: Retrospective Study
1 Pith paper cite this work, alongside 12 external citations. Polarity classification is still indexing.
1
Pith paper citing it
12
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
A new German open-response medical benchmark shows LLM judges can match physician agreement (κ≈0.69 vs 0.71) yet lack clinical caution and show lineage-dependent scoring bias.