A dialogue-based medical benchmark with a D.O.T.S. score (Diagnosis, Observations, Treatment, Steps) shows that a structured agentic system beats a bare LLM prompt and approaches human GP accuracy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI
A dialogue-based medical benchmark with a D.O.T.S. score (Diagnosis, Observations, Treatment, Steps) shows that a structured agentic system beats a bare LLM prompt and approaches human GP accuracy.