Pith. sign in

REVIEW 2 cited by

Large Language Models Perform Diagnostic Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.08922 v1 pith:5SRZDBSG submitted 2023-07-18 cs.CL

classification cs.CL
keywords reasoninglanguagelargemodelspromptingdiagnosticdr-cotaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore the extension of chain-of-thought (CoT) prompting to medical reasoning for the task of automatic diagnosis. Motivated by doctors' underlying reasoning process, we present Diagnostic-Reasoning CoT (DR-CoT). Empirical results demonstrate that by simply prompting large language models trained only on general text corpus with two DR-CoT exemplars, the diagnostic accuracy improves by 15% comparing to standard prompting. Moreover, the gap reaches a pronounced 18% in out-domain settings. Our findings suggest expert-knowledge reasoning in large language models can be elicited through proper promptings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Mid-stream alignment on next-step clinical decisions (MedUPS) improves LLM next-step accuracy on uncommon cases, but the size of the gain depends substantially on which LLM judge does the scoring.

  2. ProEvent: An Event-centric Benchmark for Proactive Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    ProEvent is a benchmark showing LLM agents keep a user's event timetable from chats poorly, with the best fully-correct score at 27.2%.

Pith tools