Pith. sign in

REVIEW 2 cited by

Onco-Retriever: Generative Classifier for Retrieval of EHR Records in Oncology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06680 v1 pith:GXLQHCW3 submitted 2024-04-10 cs.CL

classification cs.CL
keywords datainformationmodelsretrievingsystemsbetterclinicalcreating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Retrieving information from EHR systems is essential for answering specific questions about patient journeys and improving the delivery of clinical care. Despite this fact, most EHR systems still rely on keyword-based searches. With the advent of generative large language models (LLMs), retrieving information can lead to better search and summarization capabilities. Such retrievers can also feed Retrieval-augmented generation (RAG) pipelines to answer any query. However, the task of retrieving information from EHR real-world clinical data contained within EHR systems in order to solve several downstream use cases is challenging due to the difficulty in creating query-document support pairs. We provide a blueprint for creating such datasets in an affordable manner using large language models. Our method results in a retriever that is 30-50 F-1 points better than propriety counterparts such as Ada and Mistral for oncology data elements. We further compare our model, called Onco-Retriever, against fine-tuned PubMedBERT model as well. We conduct an extensive manual evaluation on real-world EHR data along with latency analysis of the different models and provide a path forward for healthcare organizations to build domain-specific retrievers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment

    cs.IR 2025-02 conditional novelty 6.0 of 10

    CliniQ is a public EHR retrieval benchmark with 77,206 LLM-annotated relevance judgments, showing that BM25 is a strong baseline and that semantic matches drive dense-retriever gains.

  2. DR.EHR: Dense Retrieval for Electronic Health Record with Knowledge Injection and Synthetic Data

    cs.IR 2025-07 conditional novelty 5.0 of 10

    DR.EHR, a two-stage trained dense retriever using BIOS knowledge injection and Llama-generated synthetic data, achieves state-of-the-art results on the CliniQ EHR retrieval benchmark.

Pith tools