Pith. sign in

REVIEW 1 cited by

Large language models are good medical coders, if provided with tools

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12849 v1 pith:V6DC3VQY submitted 2024-07-06 cs.IR cs.CL

classification cs.IRcs.CL
keywords medicalaccuracyretrieve-ranksystemachievedcodingicd-10-cmlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study presents a novel two-stage Retrieve-Rank system for automated ICD-10-CM medical coding, comparing its performance against a Vanilla Large Language Model (LLM) approach. Evaluating both systems on a dataset of 100 single-term medical conditions, the Retrieve-Rank system achieved 100% accuracy in predicting correct ICD-10-CM codes, significantly outperforming the Vanilla LLM (GPT-3.5-turbo), which achieved only 6% accuracy. Our analysis demonstrates the Retrieve-Rank system's superior precision in handling various medical terms across different specialties. While these results are promising, we acknowledge the limitations of using simplified inputs and the need for further testing on more complex, realistic medical cases. This research contributes to the ongoing effort to improve the efficiency and accuracy of medical coding, highlighting the importance of retrieval-based approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The NordDRG AI Benchmark for Large Language Models

    cs.AI 2025-06 conditional novelty 7.0 of 10

    The paper releases the first public, rule-complete benchmark for LLM reasoning over NordDRG hospital payment logic, with top models scoring 13/13 on logic tasks and 7/13 on full grouper emulation.

Pith tools