Pith. sign in

REVIEW 7 cited by

AI-based Clinical Decision Support for Primary Care: A Real-World Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.16947 v1 pith:5BZKISFN submitted 2025-07-22 cs.CL

AI-based Clinical Decision Support for Primary Care: A Real-World Study

classification cs.CL
keywords consulterrorsclinicalclinicianscarevisitscliniciandecision
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was "substantial". These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

    cs.AI 2026-05 conditional novelty 8.0

    PhysicianBench is a new benchmark of 100 physician-reviewed, execution-grounded tasks in live EHR environments where the best LLM agent reaches only 46% success and open-source models reach 19%.

  2. Towards Conversational Medical AI with Eyes, Ears and a Voice

    cs.AI 2026-05 conditional novelty 6.0

    AI co-clinician is a multimodal conversational AI that uses live audio-visual data for real-time medical reasoning in simulated telemedicine, approaching primary care physicians in management plans and differentials b...

  3. Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

    cs.AI 2026-04 unverdicted novelty 6.0

    Case-specific clinician rubrics for clinical AI notes achieve strong discrimination between outputs, high stability, and clinician-LLM agreement matching clinician-clinician levels at far lower cost.

  4. Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight

    cs.AI 2025-12 conditional novelty 6.0

    Physician oversight reveals high error rates in LLM-generated labels for a clinical benchmark and demonstrates that corrected labels improve both evaluation accuracy and downstream model training.

  5. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  6. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

    cs.AI 2026-07 accept novelty 5.0

    Passing medical exams does not make LLMs safe for autonomous triage: they fail to seek missing red flags, and current benchmarks do not test that.

  7. Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

    cs.CL 2025-12 conditional novelty 5.0

    Fine-tuned on-device LLMs achieve up to 87.9% diagnostic accuracy on clinical tasks, approaching GPT-5.1 at 89.4% while remaining smaller and local.