REVIEW 2 major objections 4 minor 81 references
CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment
T0 review · 2 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper introduces CliniQ, a public benchmark for electronic health record retrieval built on MIMIC-III with 1,246 queries, 77,206 labeled relevance judgments, and a five-way match-type taxonomy that exposes the semantic gap.
desk verdict CliniQ is a genuinely useful new EHR retrieval resource, but the Multi-Patient half runs on an incomplete gold standard the paper itself concedes; the Single-Patient half is the solid part. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the CliniQ test collection and its annotation protocol. Discharge summaries are split into 100-word overlapping chunks; queries are obtained by aggregating ICD-9 disease codes (mapped to three-digit ancestors), ICD-9 procedure codes, and cleaned prescription names; exact string search finds literal matches; and GPT-4o with a chain-of-thought prompt classifies every remaining query–chunk pair as irrelevant or as one of four semantic match types. The two retrieval settings (single-patient and multi-patient) define the evaluation regimes, and the match-type taxonomy turns the semantic gap into five concrete categories. A human-evaluated subset of 141 chunks with 5,221 annotations shows GPT-4o in near-perfect agreement with two senior medical annotators, which is the evidentiary basis for trusting the automated labels.
What would settle it
Re-annotate a random sample of patient-note chunks against the full global query set of 1,246 queries instead of only the queries assigned to those patients. If a non-trivial fraction of previously unlabeled pairs turns out to be relevant, the multi-patient recall figures are inflated and the completeness assumption fails. A second check would be to collect a fresh human-annotated holdout beyond the 141 chunks in the paper and test whether GPT-4o's match-type labels maintain the reported agreement.
Extended reading notes
Core claim
The central claim is that a large-scale public EHR retrieval benchmark can be built by combining a public corpus with LLM-based annotation, and that such a benchmark changes what can be seen about retrieval systems. CliniQ is that benchmark: 1,000 MIMIC-III discharge summaries chunked into 16,550 segments, queried with 1,246 coded entities, and labeled with 77,206 relevance judgments plus a five-way match-type taxonomy. Using it, the paper shows that BM25 is a strong baseline, that dense retrievers' advantage over BM25 is almost entirely due to semantic matches, that implication matches are the hardest category for every tested method, and that general-domain dense retrievers outperform those trained for the biomedical domain. The paper also validates the annotation pipeline against human experts, reporting near-perfect agreement on both relevance and match type.
Load-bearing premise
The load-bearing premise is that the relevance judgments are complete enough to serve as ground truth: for each patient's notes, only the queries originally coded for that patient were annotated, so any chunk that is actually relevant to a different query is silently treated as irrelevant in the multi-patient evaluation.
Editorial extensions
If this is right
- BM25 with UMLS query expansion should remain the reference baseline in future EHR retrieval studies, since it stays competitive with dense retrievers, especially in multi-patient settings.
- Dense retrieval research for EHRs should focus on semantic match types, because that is where dense models earn their advantage; implication match is the hardest category for every method tested.
- General-domain dense retrievers outperforming biomedical ones indicates that EHR text is out of distribution for current biomedical embeddings, motivating EHR-specific training data.
- RRF-style fusion of sparse and dense retrieval gives large gains on both shallow and deep metrics, making hybrid retrieval a practical direction for clinical deployment.
- Drug queries, which are short and mostly literal, favor lexical matching; dense retrievers need specific work to preserve fine-grained string information.
Reading between the lines
- The per-match-type breakdown could serve as a diagnostic suite for future systems: a model that improves implication matching without sacrificing string matching would be identifiable from CliniQ's labeled subsets, a finer signal than a single pooled score.
- The annotation recipe — exact matching plus LLM refinement over a public coded corpus — generalizes to other note types and coding systems, although CliniQ itself only covers discharge summaries.
- If adopted, CliniQ's match-type labels could be used directly as training signal, e.g., as targets for contrastive learning that rewards semantic equivalence rather than only surface similarity, which the paper does not explore.
- The benchmark's reliance on MIMIC-III access means its long-term availability depends on that dataset remaining accessible; porting the pipeline to MIMIC-IV would test the collection method's reproducibility.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CliniQ, a publicly released EHR retrieval benchmark built from 1,000 MIMIC-III discharge summaries, chunked into 16,550 passages. Queries are derived from ICD-9 disease codes, ICD-9 procedure codes, and prescription labels, yielding 1,246 unique queries. Relevance is annotated at chunk level using exact string matching plus GPT-4o judgments, with match types classified as string, synonym, abbreviation, hyponym, or implication; the full dataset contains 77,206 positive relevance judgments. The benchmark supports two settings: Single-Patient Retrieval and Multi-Patient Retrieval. The authors evaluate BM25, UMLS-based query expansion, several open-source dense retrievers, a proprietary embedding model, and reciprocal rank fusion, and report that BM25 is a strong baseline, general-domain dense retrievers outperform biomedical-domain ones, implication matches are the hardest, and fusion improves results.
Significance. If the annotation quality and completeness claims hold, CliniQ is a valuable community resource: it is, to my knowledge, the first large-scale public EHR retrieval benchmark with fine-grained semantic match categories, and it includes a human evaluation on 5,221 judgments showing very high agreement with GPT-4o. The release of code and patient identifiers for corpus reconstruction is a practical and appropriate way to handle MIMIC data-use restrictions. The benchmark also provides a broad, reproducible comparison of sparse and dense retrievers, which is useful for the community. However, the Multi-Patient relevance judgments are incomplete by construction, and this directly affects the headline Multi-Patient rankings and the claim that BM25 is a strong baseline in that setting. The Single-Patient portion and the semantic match analysis are less affected, but the Multi-Patient results need substantially more support before the benchmark can be relied upon as a fair evaluation of the two settings.
major comments (2)
- [Section 3.3, Section 4.1, Table 3] The Multi-Patient gold standard is incomplete in a way that is not neutral across methods. Section 3.3 states that semantic-match judgments are produced only for queries assigned to a patient in the coding systems Q_i, rather than for all queries Q, and Section 4.1 defines Multi-Patient Retrieval as retrieving for every query over all 16,550 chunks. Therefore, for a query q, chunks from patients whose structured codes do not include q are never judged for semantic relevance and are silently treated as irrelevant in Table 3. MIMIC ICD and prescription codes are billing-derived and incomplete, as the paper itself notes in the Limitations, so a documented condition in a non-coded patient's note can easily be missed. This creates an unquantified false-negative rate that is likely higher for dense retrievers, which surface semantically related text, than for BM25, which retrieves lexically matching chunks from coded patients. Because the Multi-Patient MRR, NDCG@10, and recall@100 in Table 3 are computed against this partial gold standard, the central conclusion that BM25 is a strong Multi-Patient baseline relative to dense retrievers is not established. The defense in the Limitations that all retrieval benchmarks contain false negatives is not fully persuasive: TREC-style collections use pooling over many systems, whereas here negative labels are generated from a single incomplete billing-code source without pooling. Please either (a) construct the Multi-Patient judgments by pooling candidate chunks from multiple retrieval systems and annotating them, (b) re-annotate a random sample of non-assigned patient-query pairs to estimate the false-negative bias and show that model rankings are robust to it, or (c) explicitly restrict the Multi-Patient claims to the annotated patient-query pairs and avoid presenting the current scores as complete relevance-based metrics.
- [Section 5.1, Table 1] Table 1 reports 77.3 relevant judgments per patient in Single-Patient Retrieval, which projects to about 77.3k over 1,000 notes, while the Multi-Patient total is 77,206. The Single-Patient match-type subcounts (29.2, 24.8, 3.0, 4.3, 15.9) are exactly the Multi-Patient totals (29,149, 24,798, 3,039, 4,288, 15,932) divided by 1,000. This is surprising because Section 4.1 says the Single-Patient query set is filtered to queries with at least one positive chunk in that note, whereas Section 3.3's global string matching should add extra string-match positives in the Multi-Patient setting beyond the within-patient positives. The two totals therefore cannot coincide in the way the table suggests. Please clarify how the Single-Patient column was computed, or correct the table; as written, it conflates the two settings and undermines confidence in the reported dataset statistics.
minor comments (4)
- [Section 5.2] There is a typo: 'Among dense retrivers with a parameter size less than 7B' should read 'Among dense retrievers with fewer than 7B parameters'.
- [Section 5.4] The heading 'Query type assessment' refers to 'Simple-Patient Retrieval' in the body text; this should be 'Single-Patient Retrieval'.
- [Section 2.1.3] The phrase 'Cohen's Kappa efficient of 1' should be 'Cohen's kappa coefficient of 1'.
- [Figure 1] The caption reports averages of MRR, NDCG, and MAP for Single-Patient Retrieval and averages of MRR, NDCG@10, and recall@100 for Multi-Patient Retrieval; since these metrics have different scales, the figure should state whether the components were normalized before averaging.
Circularity Check
No significant circularity: queries come from MIMIC structured codes, labels from GPT-4o with human agreement checks, and all evaluated retrievers are zero-shot off-the-shelf systems.
full rationale
CliniQ's construction chain is not circular. Queries are derived from MIMIC-III ICD-9 disease/procedure codes and prescription labels after cleaning (Section 3.2), not from the retrieval models or from the benchmark's own conclusions. Relevance judgments are produced by exact-match search plus GPT-4o annotation with a detailed prompt (Section 3.3, Figure 3), and the paper reports independent human-expert agreement (Table 2), so the labels are externally checked rather than self-justified. All baselines are used zero-shot with standard hyperparameters (BM25 k1=1.5, b=0.75; RRF k=60 by convention), and no model is trained or tuned on CliniQ, so no fitted parameter is renamed as a prediction. The paper's own self-citations appear in related-work and background contexts, not as load-bearing justification for the benchmark's validity. The most serious weakness—the Multi-Patient gold standard is incomplete because 'we only annotate queries assigned to the patient in the coding systems Q_i, rather than all queries Q' (Section 3.3)—is a completeness/bias limitation, and the paper explicitly acknowledges 'the incompleteness of ICD labels introduces unavoidable false negatives' in the Limitations. This can affect the fairness of model comparisons, but it is not a circular derivation: the reported metrics are not equivalent by construction to the inputs, and the benchmark's central contribution (a released query/chunk/judgment resource) does not collapse into its own assumptions.
Assumptions & free parameters
assumptions (4)
- domain assumption MIMIC-III discharge summaries and ICD/prescription codes are representative of real-world EHR retrieval queries.
- domain assumption Exact string match is a sufficient and correct way to label all string-match relevant chunks.
- domain assumption GPT-4o annotations agree with medical experts beyond the 141-chunk evaluation sample.
- domain assumption Unannotated chunks in the multi-patient setting are irrelevant.
Cite this review
Pith. "Pith review of CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment." pith.science (2026). https://pith.science/paper/MKE4ZF3A
@misc{pith2026250206252,
author = {Pith},
title = {Pith review of: CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment},
year = {2026},
howpublished = {\url{https://pith.science/paper/MKE4ZF3A}},
note = {Machine review of arXiv:2502.06252}
}
read the original abstract
Electronic Health Record (EHR) retrieval plays a pivotal role in various clinical tasks, but its development has been severely impeded by the lack of publicly available benchmarks. In this paper, we introduce a novel public EHR retrieval benchmark, CliniQ, to address this gap. We consider two retrieval settings: Single-Patient Retrieval and Multi-Patient Retrieval, reflecting various real-world scenarios. Single-Patient Retrieval focuses on finding relevant parts within a patient note, while Multi-Patient Retrieval involves retrieving EHRs from multiple patients. We build our benchmark upon 1,000 discharge summary notes along with the ICD codes and prescription labels from MIMIC-III, and collect 1,246 unique queries with 77,206 relevance judgments by further leveraging powerful LLMs as annotators. Additionally, we include a novel assessment of the semantic gap issue in EHR retrieval by categorizing matching types into string match and four types of semantic matches. On our proposed benchmark, we conduct a comprehensive evaluation of various retrieval methods, ranging from conventional exact match to popular dense retrievers. Our experiments find that BM25 sets a strong baseline and performs competitively to the dense retrievers, and general domain dense retrievers surprisingly outperform those designed for the medical domain. In-depth analyses on various matching types reveal the strengths and drawbacks of different methods, enlightening the potential for targeted improvement. We believe that our benchmark will stimulate the research communities to advance EHR retrieval systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Qwen2 Technical Report
2024. Qwen2 Technical Report. (2024)
2024
-
[2]
Negar Arabzadeh, Xinyi Yan, and Charles L. A. Clarke. 2021. Predicting Effi- ciency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection. Proceedings of the 30th ACM International Conference on Information & Knowledge Management (2021). https://api.semanticscholar.org/CorpusID:237593050
work page 2021
-
[3]
Brian G Arndt, John W Beasley, Michelle D Watkinson, Jonathan L Temte, Wen- Jan Tuan, Christine A Sinsky, and Valerie J Gilchrist. 2017. Tethered to the EHR: primary care physician workload assessment using EHR event log data and time-motion observations. The Annals of Family Medicine 15, 5 (2017), 419–426
work page 2017
-
[4]
Olivier Bodenreider. 2004. The Unified Medical Language System (UMLS): inte- grating biomedical terminology. Nucleic acids research 32 Database issue (2004), D267–70. https://api.semanticscholar.org/CorpusID:205228801
work page 2004
-
[5]
Guillaume Bouzillé, Marie-Noëlle Osmont, Louise Triquet, Natalia Grabar, Cécile Rochefort-Morel, Emmanuel Chazard, Elisabeth Polard, and Marc Cuggia. 2018. Drug safety and big clinical data: Detection of drug-induced anaphylactic shock events. Journal of Evaluation in Clinical Practice 24, 3 (2018), 536–544
work page 2018
-
[6]
Gordon V. Cormack, Charles L. A. Clarke, and Stefan Büttcher. 2009. Recipro- cal rank fusion outperforms condorcet and individual rank learning methods. Proceedings of the 32nd international ACM SIGIR conference on Research and devel- opment in information retrieval (2009). https://api.semanticscholar.org/CorpusID: 12408211
work page 2009
-
[7]
Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
arXiv 2018
-
[8]
Tracy Edinger, Aaron M Cohen, Steven Bedrick, Kyle Ambert, and William Hersh
Show all 81 references
-
[9]
Shashi Kant Gupta, Aditya Basu, Bradley Taylor, Anai Kothari, and Hrituraj Singh. 2024. Onco-Retriever: Generative Classifier for Retrieval of EHR Records in Oncology. arXiv:2404.06680 [cs.CL] https://arxiv.org/abs/2404.06680
2024 arXiv
-
[10]
Kenric W Hammond, Ryan J Laundry, T Michael OLeary, and William P Jones
-
[11]
Hanauer, Qiaozhu Mei, James Law, Ritu Khanna, and Kai Zheng
David A. Hanauer, Qiaozhu Mei, James Law, Ritu Khanna, and Kai Zheng. 2015. Supporting information retrieval from electronic health records: A report of Uni- versity of Michigan’s nine-year experience in developing and using the Electronic Medical Record Search Engine (EMERSE)...
2015
-
[12]
Rachael Hopkins. 2004. Information retrieval: a health and biomed- ical perspective. Health Information & Libraries Journal 21, 4 (2004), 277–278. https://doi.org/10.1111/j.1471-1842.2004.00530.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1471-1842.2004.00530.x
2004
-
[13]
Kasra Hosseini, Thomas Kober, Josip Krapac, Roland Vollgraf, Weiwei Cheng, and Ana Peleteiro-Ramallo. 2024. Retrieve, Annotate, Evaluate, Repeat: Lever- aging Multimodal LLMs for Large-Scale Product Retrieval Evaluation. ArXiv abs/2409.11860 (2024). https://api.semanticscholar...
2024 arXiv
-
[14]
Richard G. Jackson, Ismail Emre Kartoglu, Clive Stringer, Genevieve Gorrell, Angus Roberts, Xingyi Song, Honghan Wu, Asha Agrawal, Kenneth Lui, Tudor Groza, Damian Lewsley, Doug Northwood, Amos A. Folarin, Robert J Stewart, and Richard J. B. Dobson. 2017. CogStack - experience...
2017
-
[15]
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...
2023 arXiv
-
[16]
Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. MedCPT: Contrastive Pre-trained Transform- ers with large-scale PubMed search logs for zero-shot biomedical information retrieval. Bioinformatics 39, 11 (2023), btad651
2023
-
[17]
Qiao Jin, Chuanqi Tan, Zhengyun Zhao, Zheng Yuan, and Songfang Huang. 2021. Alibaba DAMO Academy at TREC Clinical Trials 2021: ExploringEmbedding- based First-stage Retrieval with TrialMatcher. In Text Retrieval Conference. https: //api.semanticscholar.org/CorpusID:247849810
2021
-
[18]
Floudas, Jimeng Sun, and Zhiyong Lu
Qiao Jin, Zifeng Wang, Charalampos S. Floudas, Jimeng Sun, and Zhiyong Lu
-
[19]
Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al
-
[20]
Alistair E. W. Johnson, Tom J. Pollard, Lu Shen, Li wei H. Lehman, Mengling Feng, Mohammad Mahdi Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. 2016. MIMIC-III, a freely accessible critical care database. Scientific Data 3 (2016). https://api.s...
2016
-
[21]
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open- domain question answering. arXiv preprint arXiv:2004.04906 (2020)
2020 arXiv
-
[22]
Ismail Mohamed Keshta and Ammar Jamil Odeh. 2020. Security and privacy of electronic health records: Concerns and challenges. Egyptian Informatics Journal (2020). https://api.semanticscholar.org/CorpusID:225426124
2020
-
[23]
Scientific data 10, 1 (2023), 1
MIMIC-IV, a freely accessible electronic health record dataset. Scientific data 10, 1 (2023), 1
2023
-
[24]
Vojtech Lanz and Pavel Pecina. 2024. Paragraph Retrieval for Enhanced Question Answering in Clinical Documents. In Workshop on Biomedical Natural Language Processing. https://api.semanticscholar.org/CorpusID:271769434
2024
-
[25]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv preprint arXiv:2405.17428 (2024)
2024 arXiv
-
[26]
Mengyang Li, Hailing Cai, Shan Nan, Jialin Li, Xudong Lu, and Huilong Duan
-
[27]
Bevan Koopman, Guido Zuccon, Peter Bruza, Laurianne Sitbon, and Michael Lawley. 2016. Information retrieval as semantic inference: A graph inference model applied to medical search. Information Retrieval Journal 19 (2016), 6–37
2016
-
[28]
Hersh, and Hongfang Liu
Sijia Liu, Yanshan Wang, Andrew Wen, Liwei Wang, Na Hong, Feichen Shen, Steven Bedrick, William R. Hersh, and Hongfang Liu. 2020. Implementation of a Cohort Retrieval System for Clinical Data Repositories Using the Observational Medical Outcomes Partnership Common Data Model: ...
2020
-
[29]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. ArXiv abs/2310.08319 (2023). https://api.semanticscholar.org/CorpusID:263908865
2023 arXiv
-
[30]
David Martinez, Arantxa Otegi, Aitor Soroa, and Eneko Agirre. 2014. Improv- ing search over Electronic Health Records using UMLS-based query expansion through random walks. Journal of biomedical informatics 51 (2014), 100–106
2014
-
[31]
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022. MTEB: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316 (2022)
2022 arXiv
-
[32]
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meis- han Zhang. 2023. Towards General Text Embeddings with Multi-stage Con- trastive Learning. ArXiv abs/2308.03281 (2023). https://api.semanticscholar.org/ CorpusID:260682258
2023 arXiv
-
[33]
Miller, Yanjun Gao, Matthew M
Skatje Myers, Timothy A. Miller, Yanjun Gao, Matthew M. Churpek, Anoop M. Mayampurath, Dmitriy Dligach, and Majid Afshar. 2024. Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies. Journal of the Americ...
2024
-
[34]
Stein, Samat Jain, and Noémie Elhadad
Karthik Natarajan, Daniel M. Stein, Samat Jain, and Noémie Elhadad. 2010. An analysis of clinical queries in an electronic health record search utility. International journal of medical informatics 79 7 (2010), 515–22. https://api. semanticscholar.org/CorpusID:18473561
2010
-
[35]
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al
-
[36]
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers. arXiv preprint arXiv:2112.07899 (2021)
2021 arXiv
-
[37]
Mullenbach, Sarah Wiegreffe, Jon D
J. Mullenbach, Sarah Wiegreffe, Jon D. Duke, Jimeng Sun, and Jacob Eisenstein
-
[38]
Liang, and Jian Peng
Anusri Pampari, Preethi Raghavan, Jennifer J. Liang, and Jian Peng. 2018. emrQA: A Large Corpus for Question Answering on Electronic Medical Records. In Conference on Empirical Methods in Natural Language Processing . https://api. semanticscholar.org/CorpusID:52158121
2018
-
[39]
Zhang Ping and Wu Jinfa. 2021. Research on Search Ranking Technology of Chinese Electronic Medical Record Based on Adarank. 2021 18th International Computer Conference on Wavelet Active Media Technology and Information Pro- cessing (ICCW AMTIP)(2021), 63–70. https://api.semant...
2021
-
[40]
Robertson and Hugo Zaragoza
Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. 3 (2009), 333–389. https: //api.semanticscholar.org/CorpusID:207178704
2009
-
[41]
Halley Ruppel, Aashish Bhardwaj, Raj N Manickam, Julia Adler-Milstein, Marc Flagg, Manuel Ballesca, and Vincent X Liu. 2020. Assessment of electronic health record search patterns and practices by practitioners in a large integrated health care system. JAMA network open 3, 3 (...
2020
-
[42]
Savova, James J
Guergana K. Savova, James J. Masanz, Philip V. Ogren, Jiaping Zheng, Sunghwan Sohn, Karin Kipper Schuler, and Christopher G. Chute. 2010. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications. Journal of the A...
2010
-
[43]
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Sim- ple Entity-Centric Questions Challenge Dense Retrievers. ArXiv abs/2109.08535 (2021). https://api.semanticscholar.org/CorpusID:237562875
2021 arXiv
-
[44]
Venet Osmani, Stefano Forti, Oscar Mayora, and Diego Conforti. 2017. Chal- lenges and opportunities in evolving TreC personal health record platform. In Proceedings of the 11th EAI International Conference on Pervasive Comput- ing Technologies for Healthcare (Barcelona, Spain)...
2017
-
[45]
Syeda-Mahmood, and Tyler Baldwin
Luyao Shi, Tanveer F. Syeda-Mahmood, and Tyler Baldwin. 2022. Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation. In North American Chapter of the Association for Computational Linguistics. https://api.semanticscholar.org/CorpusID:250390560
2022
-
[46]
Sonish Sivarajkumar, Haneef Ahamed Mohammad, David Oniani, Kirk Roberts, William Hersh, Hongfang Liu, Daqing He, Shyam Visweswaran, and Yanshan Wang. 2024. Clinical information retrieval: A literature review. Journal of Healthcare Informatics Research (2024), 1–40
2024
-
[47]
Sarvesh Soni and Kirk Roberts. 2020. Patient Cohort Retrieval using Transformer Language Models. AMIA ... Annual Symposium proceedings. AMIA Symposium 2020 (2020), 1150–1159. https://api.semanticscholar.org/CorpusID:221640646
2020
-
[48]
Bo Sun, Fei Zhang, Jing Li, Yicheng Yang, Xiaolin Diao, Wei Zhao, and Ting Shu
-
[49]
Lynda Tamine and Lorraine Goeuriot. 2021. Semantic information retrieval on medical texts: Research challenges, survey, and open issues. ACM Computing Surveys (CSUR) 54, 7 (2021), 1–38
2021
-
[50]
Dung Ngoc Thai, Victor Ardulov, Jose Ulises Mena, Simran Tiwari, Gleb Erofeev, Ramy Eskander, Karim Tarabishy, Ravi B Parikh, and Wael Salloum. 2024. ACR: Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Zhengyun Zhao, Hongyi Yuan, Jingjing Liu, Haichao Chen, Huaiyuan Y...
2024 arXiv
-
[51]
Shivani Shekhar, Simran Tiwari, T. C. Rensink, Ramy Eskander, and Wael Sal- loum. 2023. Coupling Symbolic Reasoning with Language Modeling for Efficient Longitudinal Understanding of Unstructured Electronic Medical Records. ArXiv abs/2308.03360 (2023). https://api.semanticscho...
2023 arXiv
-
[52]
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, and Jimmy Lin. 2024. A Large- Scale Study of Relevance Assessments with Large Language Models: An Initial Look. ArXiv abs/2411.08275 (2024). https://api.semanticscholar....
2024 arXiv
-
[53]
Mayrdorfer, Klemens Budde, Felix Alexander Gers, and Alexander Loser
Betty van Aken, Jens-Michalis Papaioannou, M. Mayrdorfer, Klemens Budde, Felix Alexander Gers, and Alexander Loser. 2021. Clinical Outcome Prediction from Admission Notes using Self-Supervised Knowledge Integration. In Con- ference of the European Chapter of the Association fo...
2021
-
[54]
Voorhees
Ellen M. Voorhees. 2013. The TREC Medical Records Track. In Proceedings of the International Conference on Bioinformatics, Computational Biology and Biomedical Informatics (Wshington DC, USA) (BCB’13). Association for Computing Machin- ery, New York, NY, USA, 239–246. https://...
2013
-
[55]
Hersh, Steven Bedrick, and Hongfang Liu
Yanshan Wang, Andrew Wen, Sijia Liu, William R. Hersh, Steven Bedrick, and Hongfang Liu. 2019. Test collections for electronic health record-based clinical in- formation retrieval. JAMIA Open 2 (2019), 360 – 368. https://api.semanticscholar. org/CorpusID:195441753
2019
-
[56]
BMC Medical Informatics and Decision Making 21 (2021)
Using NLP in openEHR archetypes retrieval to promote interoperability: a feasibility study in China. BMC Medical Informatics and Decision Making 21 (2021). https://api.semanticscholar.org/CorpusID:235638693
2021
-
[57]
Cao Xiao, Junyi Gao, Lucas Glass, and Jimeng Sun. 2020. Patient trial matching using pseudo-siamese network. Journal of Clinical Oncology 38 (2020). https: //api.semanticscholar.org/CorpusID:219780496
2020
-
[58]
Shitao Xiao, Zheng Liu, Yingxia Shao, and Zhao Cao. 2022. RetroMAE: Pre- Training Retrieval-oriented Language Models Via Masked Auto-Encoder. In Conference on Empirical Methods in Natural Language Processing . https://api. semanticscholar.org/CorpusID:252917569
2022
-
[59]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663 (2021)
2021 arXiv
-
[60]
Ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang, Yanqiao Zhu, May D Wang, Joyce C Ho, Chao Zhang, and Carl Yang. 2024. Bmretriever: Tuning large language models as better biomedical text retrievers. arXiv preprint arXiv:2404.18443 (2024)
2024 arXiv
-
[61]
Lei Yang, Qiaozhu Mei, Kai Zheng, and David A. Hanauer. 2011. Query log analysis of an electronic health record search engine.AMIA ... Annual Symposium proceedings. AMIA Symposium 2011 (2011), 915–24. https://api.semanticscholar. org/CorpusID:35582937
2011
-
[62]
Songchun Yang, Xiangwen Zheng, Yu Xiao, Xiangfei Yin, Jianfei Pang, Huajian Mao, Wei Wei, Wenqin Zhang, Yu Yang, Haifeng Xu, Mei Li, and Dongsheng Zhao. 2021. Improving Chinese electronic medical record retrieval by field weight assignment, negation detection, and re-ranking. ...
2021
-
[63]
Cheng Ye and Daniel Fabbri. 2018. Extracting similar terms from multiple EMR- based semantic embeddings to support chart reviews. Journal of biomedical informatics 83 (2018), 63–72
2018
-
[64]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models.ArXiv abs/2201.11903 (2022). https://api.semanticscholar. org/CorpusID:246411621
2022 arXiv
-
[65]
Huaiyuan Ying, Hongyi Yuan, Jinsen Lu, Zitian Qu, Yang Zhao, Zhengyun Zhao, Isaac Kohane, Tianxi Cai, and Sheng Yu. 2025. GENIE: Generative Note Infor- mation Extraction model for structuring EHR data. arXiv:2501.18435 [cs.CL] https://arxiv.org/abs/2501.18435
2025 arXiv
-
[66]
Huaiyuan Ying, Zhengyun Zhao, Yang Zhao, Sihang Zeng, and Sheng Yu. 2024. CoRTEx: contrastive learning for representing terms via explanations with appli- cations on constructing biomedical knowledge graphs. Journal of the American Medical Informatics Association (2024), ocae115
2024
-
[67]
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv:2309.07597 [cs.CL]
2023 arXiv
-
[68]
Lu, Jing Wang, Yutao Xie, and Heung yeung Shum
Sheng Yu, Zheng Yuan, Jun Xia, Shengxuan Luo, Huaiyuan Ying, Sihang Zeng, Jingyi Ren, Hongyi Yuan, Zhengyun Zhao, Yucong Lin, K. Lu, Jing Wang, Yutao Xie, and Heung yeung Shum. 2022. BIOS: An Algorithmically Gen- erated Biomedical Knowledge Graph. ArXiv abs/2203.09975 (2022). ...
2022 arXiv
-
[69]
Zheng Yuan, Chuanqi Tan, and Songfang Huang. 2022. Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding. In Annual Meeting of the Association for Computational Linguistics . https://api. semanticscholar.org/CorpusID:247222710
2022
-
[70]
Zheng Yuan, Zhengyun Zhao, and Sheng Yu. 2020. CODER: Knowledge-infused cross-lingual medical term embedding for term normalization. Journal of biomedical informatics (2020), 103983. https://api.semanticscholar.org/CorpusID: 226254376
2020
-
[71]
Yichi Zhang, Tianrun Cai, Sheng Yu, Kelly Cho, Chuan Hong, Jiehuan Sun, Jie Huang, Yuk-Lam Ho, Ashwin N Ananthakrishnan, Zongqi Xia, et al. 2019. High- throughput phenotyping with electronic medical record data using a common semi-supervised approach (PheCAP). Nature protocols...
2019
-
[72]
Cheng Ye, Bradley A Malin, and Daniel Fabbri. 2021. Leveraging medical con- text to recommend semantically similar terms for chart reviews. BMC Medical Informatics and Decision Making 21, 1 (2021), 353
2021
-
[73]
Zuccon, and Daxin Jiang
Shengyao Zhuang, Linjun Shou, Jian Pei, Ming Gong, Houxing Ren, G. Zuccon, and Daxin Jiang. 2023. Typos-aware Bottlenecked Pre-Training for Robust Dense Retrieval. Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval...
2023
-
[75]
Ravi, Peter Brune, Fanyu Kong, Dave Anderson, George Lee, Arie Meir, Farhana Bandukwala, Elli Kanal, Sercan Ö
Jinsung Yoon, Michel Mizrahi, Nahid Farhady Ghalaty, Thomas Dunn Jarvinen, Ashwin S. Ravi, Peter Brune, Fanyu Kong, Dave Anderson, George Lee, Arie Meir, Farhana Bandukwala, Elli Kanal, Sercan Ö. Arik, and Tomas Pfister. 2023. EHR- Safe: generating high-fidelity and privacy-pr...
2023
-
[80]
Zhengyun Zhao, Qiao Jin, Fangyuan Chen, Tuorui Peng, and Sheng Yu. 2023. A large-scale dataset of patient summaries for retrieval-based clinical decision support systems. Scientific Data 10 (12 2023). https://doi.org/10.1038/s41597- 023-02814-8
2023 doi
-
[2012]
In AMIA annual symposium proceedings, Vol
Barriers to retrieving patient information from electronic health record data: failure analysis from the TREC medical records track. In AMIA annual symposium proceedings, Vol. 2012. American Medical Informatics Association, 180
2012
-
[2013]
In 2013 46th Hawaii International Conference on System Sciences
Use of text search to effectively identify lifetime prevalence of suicide attempts among veterans. In 2013 46th Hawaii International Conference on System Sciences. IEEE, 2676–2683
2013
-
[2018]
In North American Chapter of the Association for Computational Linguistics
Explainable Prediction of Medical Codes from Clinical Text. In North American Chapter of the Association for Computational Linguistics . https://api. semanticscholar.org/CorpusID:3305987
-
[2021]
JMIR Medical Informatics 9, 10 (2021), e33192
A patient-screening tool for clinical research based on electronic health records using OpenEHR: development study. JMIR Medical Informatics 9, 10 (2021), e33192
2021
-
[2022]
arXiv preprint arXiv:2201.10005 (2022)
Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005 (2022)
2022 arXiv
-
[2023]
ArXiv (2023)
Matching Patients to Clinical Trials with Large Language Models. ArXiv (2023). https://api.semanticscholar.org/CorpusID:260203054
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.