In a physician-rated crowdsourced study, 76% of LLM responses to everyday health queries were valid, with GPT-4o highest (85%) and Llama3-8b lowest (50%); RAG did not consistently improve responses.
R.; Mahajan, S.; Chaurasia, A.; et al
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases
In a physician-rated crowdsourced study, 76% of LLM responses to everyday health queries were valid, with GPT-4o highest (85%) and Llama3-8b lowest (50%); RAG did not consistently improve responses.