Pith. sign in

REVIEW 4 cited by

A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.04709 v1 pith:26L2DMNN submitted 2023-08-09 cs.CL

classification cs.CL
keywords llmsmodelsmedicalmultiple-choiceabilityclaudefieldgpt-4
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, there have been significant breakthroughs in the field of natural language processing, particularly with the development of large language models (LLMs). These LLMs have showcased remarkable capabilities on various benchmarks. In the healthcare field, the exact role LLMs and other future AI models will play remains unclear. There is a potential for these models in the future to be used as part of adaptive physician training, medical co-pilot applications, and digital patient interaction scenarios. The ability of AI models to participate in medical training and patient care will depend in part on their mastery of the knowledge content of specific medical fields. This study investigated the medical knowledge capability of LLMs, specifically in the context of internal medicine subspecialty multiple-choice test-taking ability. We compared the performance of several open-source LLMs (Koala 7B, Falcon 7B, Stable-Vicuna 13B, and Orca Mini 13B), to GPT-4 and Claude 2 on multiple-choice questions in the field of Nephrology. Nephrology was chosen as an example of a particularly conceptually complex subspecialty field within internal medicine. The study was conducted to evaluate the ability of LLM models to provide correct answers to nephSAP (Nephrology Self-Assessment Program) multiple-choice questions. The overall success of open-sourced LLMs in answering the 858 nephSAP multiple-choice questions correctly was 17.1% - 25.5%. In contrast, Claude 2 answered 54.4% of the questions correctly, whereas GPT-4 achieved a score of 73.3%. We show that current widely used open-sourced LLMs do poorly in their ability for zero-shot reasoning when compared to GPT-4 and Claude 2. The findings of this study potentially have significant implications for the future of subspecialty medical training and patient care.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NOVO: Unlearning-Compliant Vision Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    NOVO is a vision transformer that forgets classes at inference time by removing learned class keys, trained with simulated unlearning to generalize to any forget set.

  2. From Image Captioning to Visual Storytelling

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

  3. GraphTrafficGPT: Enhancing Traffic Management Through Graph-Based AI Agent Coordination

    cs.AI 2025-07 reject novelty 4.0 of 10

    GraphTrafficGPT replaces TrafficGPT's sequential task chain with a graph-based agent scheduler, reporting 50.2% lower token use, 19.0% lower latency, and parallel multi-query handling.

  4. LLMs-Healthcare : Current Applications and Challenges of Large Language Models in various Medical Specialties

    cs.CL 2023-10 unverdicted novelty 2.0 of 10

    A review summarizing LLM applications for diagnostics and treatment in oncology, dermatology, dentistry, neurodegenerative disorders, and mental health, plus integration challenges.

Pith tools