{"id":"56e17f0e-4f98-4265-a78b-1afaf44789dd","arxiv_id":"2506.06340","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A position paper surveying language model approaches to EHR decision support, with illustrative, not real, experimental results.","lead":"This paper argues that language models can pull useful meaning from doctors' free-text notes in electronic health records, and that combining those notes with medical codes improves predictions. It is a position and review paper, not a study with real experiments, so the presented results are explicitly hypothetical.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'we demonstrate' is unsupported: all Section 4 results are explicitly hypothetical, and Appendix A.4's unlabeled 'ablation' table blurs the real/hypothetical boundary.","rationale":"The reader's verdict already captures the core issue: the Results section is explicitly hypothetical, so the main claim is not demonstrated. My stress-test confirms that and adds one concrete observation: the Appendix A.4 ablation (Table 5) drops the 'Hypothetical Results' label while describing a three-seed experiment, which makes the boundary between real and hypothetical results internally inconsistent within the manuscript. The correct disposition is unchanged: UNVERDICTED, because the paper offers no empirical content that can be verified, and acceptance or rejection would require evidence that is absent. I do not see a separate technical flaw in the methods exposition; the LoRA equations, attention definitions, and evaluation advice are standard and presented coherently. The problem is squarely the evidentiary status of the central claim.","tokens_in":10320,"tokens_out":5886,"duration_ms":60639,"concrete_test":"Run the protocol described in §3 and Appendix A on MIMIC-III (e.g., in-hospital mortality): fine-tune ClinicalBERT on the text notes with the stated LoRA settings, train XGBoost on the structured features, and evaluate with the §3.5 metrics. Compare the observed AUC gap to Table 2 (0.89 vs. 0.85) and the institution-transfer gap to Table 3 within the same task. If the text-based models do not reproduce at least the direction and rough magnitude of the claimed advantage, the abstract's 'demonstrate' fails in the only setting the paper names. If the run cannot be performed because the paper provides no code, splits, or trained models, that absence itself confirms the central claim is unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing problem is that the central claim — 'we demonstrate that language models ... improve predictive performance, but also generalize more effectively across institutional boundaries' — has no empirical support inside the manuscript. Section 4 opens with 'we present hypothetical results,' and Tables 2-4 carry the caption '(Hypothetical Results).' The quantitative content of the claim (e.g., Clinical ModernBERT AUC 0.91 vs. XGBoost 0.85 in Table 2; 0.83 vs. 0.72 cross-institution in Table 3) is therefore not a report of measurements, and the abstract's 'demonstrate' is contradicted by the body. The problem is not merely that the numbers might be unrepresentative; the document itself flags them as explicitly hypothetical. Appendix A.4 deepens the inconsistency: Table 5 is presented as an 'Ablation Study' with 'results ... averaged over 3 seeds,' without a hypothetical label, so the reader is invited to treat the same kind of hypothetical numbers as an actual experiment. No dataset splits, model checkpoints, or code are provided for any of these tables. The central claim is thus internally unsupported; this is a correctness risk about the manuscript's own evidentiary status, not a dispute with the field's consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that language models applied to free-text clinical notes can improve predictive performance and cross-institution generalization in EHR-based decision support, and that augmenting them with medical codes (ICD, CPT) through approaches such as Clinical ModernBERT yields further gains. It surveys relevant model families, describes standard methods for feature extraction, code integration, parameter-efficient fine-tuning with LoRA, and evaluation practices, and then presents numerical results in Section 4 that are explicitly labeled \"hypothetical results.\" The discussion covers limitations and future directions, including interpretability, generalization, fairness, and synthetic data.","tokens_in":10550,"tokens_out":7353,"duration_ms":66392,"significance":"If supported by real experiments, the claim that text-based LM features outperform structured EHR features and generalize across institutions would be practically important for clinical decision support. The paper competently summarizes relevant literature and standard modeling techniques, and it correctly recommends rigorous evaluation practices such as bootstrap confidence intervals, subgroup analysis, and cross-institutional testing. However, the manuscript contains no real empirical evaluation: all quantitative results are explicitly hypothetical, and key premises about the advantages of Clinical ModernBERT and code-description integration are drawn from the authors' own prior work. Consequently, the central empirical claims are not established in this manuscript; as a tutorial or position paper it would need to be reframed, and as an empirical study it is currently not significant.","major_comments":[{"comment":"The abstract states \"We demonstrate that language models can extract meaningful representations from unstructured notes that not only improve predictive performance, but also generalize more effectively across institutional boundaries,\" but Section 4 opens with \"we present hypothetical results\" and Tables 2-4 are captioned \"(Hypothetical Results).\" The numbers that are supposed to support the paper's central claims (e.g., Clinical ModernBERT AUC 0.91 vs. XGBoost 0.85 in Table 2; 0.83 vs. 0.72 cross-institution in Table 3) are therefore invented illustrations, not reported measurements. No dataset splits, model checkpoints, or code are provided. The word \"demonstrate\" is thus unsupported; either real experiments must be reported or the claim must be explicitly downgraded to a proposal.","section":"Section 4, Tables 2-4; Abstract"},{"comment":"Table 5 is presented as an \"Ablation Study\" with the sentence \"Results are averaged over 3 seeds\" and no \"hypothetical\" label, unlike Tables 2-4. This unlabeled table presents the same kind of fictional numbers as though they were real experimental outcomes, blurring the distinction between illustration and evidence. The table must be explicitly labeled as hypothetical or replaced with actual runs, including the variance across seeds.","section":"Appendix A.4, Table 5"},{"comment":"The limitations paragraph concedes that generalization \"can degrade significantly when applied to different health systems, even when leveraging textual features\" and that clinical narratives \"are shaped by institutional culture, billing practices, and individual provider styles.\" These statements directly contradict the abstract's claim that language models generalize \"more effectively across institutional boundaries\" and the cross-institution gains shown in Table 3. The manuscript must either provide evidence that the text-based advantage survives realistic distribution shift or soften the central claim accordingly.","section":"Section 5.1"},{"comment":"The key premises about the superiority of text-based representations and the benefit of code-description integration are drawn from the authors' own prior works (Lee et al., 2024a; Lee et al., 2025), and the paper does not independently test them. For a survey this would be acceptable, but then the abstract should not claim to \"demonstrate\" those results; the distinction between summarizing prior work and presenting new evidence must be made explicit in the text.","section":"Sections 2 and 3.3"}],"minor_comments":[{"comment":"There are placeholder citations such as \"[cite: 22]\" and \"[cite: 30]\" that should be replaced with proper references.","section":"Section 3.1"},{"comment":"The reference list has incomplete entries: \"Vaswani, A. Attention is all you need\" and \"Devlin, J. BERT: Pre-training...\" omit co-authors and should be formatted as Vaswani et al. and Devlin et al.","section":"References"},{"comment":"In the sentence about commercial-scale LLMs, \"OpenAI's 01 system\" should read \"OpenAI's O1 system.\"","section":"Section 2"},{"comment":"The contrastive loss function Lcode is mentioned but never explicitly defined; provide the equation or remove the reference to a specific loss.","section":"Section 3.3"},{"comment":"If Table 5 were real data, the paper's own evaluation framework in Section 3.5 would require confidence intervals or standard errors rather than only point estimates averaged over three seeds.","section":"Appendix A.4, Table 5"}],"recommendation":"reject","confidential_remarks":"The manuscript is essentially a survey with fabricated (though explicitly labeled) results in Section 4; the abstract's \"demonstrate\" is contradicted by the body, and Appendix A.4's unlabeled ablation table makes the evidentiary status worse. Even as a position paper, the internal contradiction between Section 5.1 and Table 3 needs resolution. The paper also leans heavily on the authors' own prior works without acknowledging this as a potential source of bias. For the claims made, the lack of real experiments is a load-bearing problem that cannot be addressed by modest revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this is not a research paper. It is a pedagogical overview of existing LM approaches to EHRs, with a Results section that openly says 'we present hypothetical results' and tables labeled '(Hypothetical Results)'. Nothing in the manuscript measures anything. There is no new method, no real experiment, no data or code. The abstract's 'we demonstrate that language models can extract meaningful representations...' is contradicted by the body of the paper.\n\nWhat the paper does well is limited but real. The mathematical background (attention, LoRA, cross-entropy) is correct and cleanly written. The related-work survey is serviceable and points to the right literature, including the authors' own Clinical ModernBERT and DK-BEHR. The Limitations section is honest about interpretability, generalization, documentation bias, and scalability. As a tutorial skeleton, the structure is fine.\n\nThe soft spots are large. First, the central evidence is invented numbers. The claim that text-based features improve performance and cross-institutional generalization is supported only by tables the authors chose to illustrate the narrative. Those numbers cannot be checked because no dataset splits, checkpoints, or code are provided. Second, the manuscript is unfinished: it uses the ICML 2025 template header, contains placeholder citations like '[cite: 12]' and '[cite: 22]', and references 'MedBERT' inconsistently. Third, Appendix A.4 is worse than Section 4: Table 5 is presented as an 'Ablation Study' with 'results averaged over 3 seeds' and no hypothetical label, so the reader is led to believe real experiments were run. That blurring is a genuine integrity problem, even in a draft.\n\nNone of this is a dispute with the field's consensus; clinical LMs and code-aware pretraining are active and promising areas. The issue is that this particular manuscript provides no evidence for its claims. It could be salvaged by resubmitting as an explicitly non-empirical position/survey paper without 'demonstrate' language, or by actually running the experiments. As submitted, it deserves a desk reject, not referee time.","headline":"An unfinished draft whose only results are explicitly hypothetical, despite an abstract claiming 'we demonstrate'; not ready for peer review.","tokens_in":11082,"tokens_out":2609,"would_cite":false,"duration_ms":25086,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that language models reading clinical notes can extract representations that beat structured EHR features and transfer across institutions, though its supporting tables are hypothetical.","keywords":["electronic health records","clinical language models","unstructured clinical notes","medical code integration","cross-institution generalization","clinical decision support","parameter-efficient fine-tuning"],"falsifier":"Run the paper's claimed comparison on a real multi-institution EHR corpus: train Clinical ModernBERT on text plus ICD codes from one hospital and XGBoost on structured features from the same hospital, then test both on a held-out hospital's data. The central claim fails if the text-based model does not beat the structured-feature baseline on the external site by a clinically meaningful margin, or if the text embedding shows the same institutional shift as the code distributions.","tokens_in":10099,"feed_emoji":"🩺","tokens_out":6956,"duration_ms":60684,"temperature":0.7,"pith_summary":"The paper sets out to establish that language models reading free-text clinical notes can extract representations that beat or complement structured electronic health record (EHR) features on diagnostic prediction, and that these text-derived features transfer better across institutions than structured codes do. The underlying mechanism it proposes is that clinical narratives carry semantic information—patient condition, treatment response, reasoning—that billing codes and lab tables compress away. To make that case, it describes a hybrid pipeline: clinical language models (BERT variants trained on medical text) such as ClinicalBERT and Clinical ModernBERT encode notes, medical codes are injected alongside their textual descriptions, and parameter-efficient fine-tuning adapts the models cheaply. A sympathetic reader would care because portable, text-based representations would attack a known bottleneck in clinical AI, namely models that fail when moved from one hospital system to another. The numerical comparisons in the paper are explicitly labelled hypothetical, so the contribution is a defended research program rather than a measured result.","feed_headline":"Can clinical notes outperform structured EHR data? This paper says yes","feed_subtitle":"Language-model features from free-text notes promise better predictions and cross-hospital portability.","key_machinery":"The load-bearing mechanism is the contextualized text embedding produced by a pretrained biomedical language model, pooled into a single vector per clinical note. Transformer self-attention computes each token's representation as a weighted sum over the note, so words are understood in context rather than as isolated features. On top of the plain text encoder, the paper adds a second mechanism—medical-code integration—where a model is trained to keep the embedding of a code close to the embedding of its textual description (for example, concatenating 'code: Ck description: D(Ck)' into the input or using a contrastive loss), which grounds the free-text space in structured ontology semantics. Parameter-efficient fine-tuning via low-rank adaptation (LoRA) supplies the practical third mechanism, making it feasible to adapt these large encoders in clinical settings. Together the text encoder, the code-description alignment, and cheap fine-tuning carry the argument that hybrid text-plus-code representations are accurate and institution-invariant.","core_discovery":"On its own terms, the paper claims that semantically informed text features are the missing ingredient in EHR machine learning: models that consume clinical notes through pretrained language models outperform models limited to structured EHR fields (Table 2), generalize to unseen institutions with a smaller performance drop (Table 3), and gain further from training that pairs medical codes with their natural-language descriptions (Table 4). It identifies Clinical ModernBERT as the strongest configuration, with a hypothetical AUC-ROC of 0.91 for diagnostic classification versus 0.85 for XGBoost on structured data, and attributes the gap to the model's joint encoding of long clinical context and coded concepts. The paper's central claim is thus that unstructured notes are not a nuisance modality to be discarded but the semantically richest view of the patient record, and that aligning text with ontologies such as ICD codes makes the representation both more accurate and more portable.","pith_inferences":["Because the reported gains are hypothetical, a fair reading is that the paper stakes out a testable priority: text-derived features will beat structured features at external sites. A direct reimplementation, which the paper does not provide, is the natural next step.","If note text truly is the harmonizing signal, the benefit should be weakest for hospitals with terse or fully template-driven notes, so the claim implies a boundary condition the paper does not explore.","Code-description alignment might generalize beyond ICD to ontologies the model never saw, since the representation is learned from language rather than from a fixed code table; this extension is not tested in the paper.","The hypothetical comparison suggests a concrete threshold a real study must beat: Clinical ModernBERT at 0.91 AUC versus XGBoost at 0.85, with the external-transfer gap similarly quantified."],"forward_implications":["If text features carry the claimed signal, clinical NLP should become a standard column of EHR risk prediction rather than an optional extra.","Cross-institution deployment could shift from per-site retraining to shared text-based encoders that harmonize documentation differences.","Pretraining recipes for clinical models should routinely pair medical codes with their textual descriptions, since the paper's tables tie code-description supervision to higher accuracy.","Evaluation of clinical models should adopt the paper's proposed norms: bootstrapped confidence intervals, subgroup-stratified reporting, and institution-disjoint splits."],"supporting_citations":[{"why":"Supplies the ClinicalBERT encoder that turns clinical notes into contextualized text features.","marker":"(Alsentzer et al., 2019)"},{"why":"Supplies Clinical ModernBERT, the paper's strongest configuration, with long-context encoding and joint code-plus-description training.","marker":"(Lee et al., 2025)"},{"why":"Demonstrates the code-description integration strategy that the paper adopts for grounding text in medical ontologies.","marker":"(An et al., 2025)"},{"why":"Provides MIMIC-III as the canonical source of real clinical notes and structured EHR fields the proposed pipeline would train on.","marker":"(Johnson et al., 2016)"},{"why":"Defines the XGBoost structured-data baseline whose 0.85 AUC the text-based models are claimed to beat.","marker":"(Chen & Guestrin, 2016)"},{"why":"Prior evidence invoked for the claim that text-based representations outperform and generalize better than traditional EHR encodings.","marker":"(Lee et al., 2024a)"},{"why":"Frames cross-institution generalization as the governing challenge that the paper's text-feature proposal is meant to address.","marker":"(Goetz et al., 2024)"},{"why":"Supplies LoRA, the parameter-efficient fine-tuning method the paper relies on to adapt clinical encoders cheaply.","marker":"(Hu et al., 2021)"}],"fun_headline_variants":["LLMs turn clinical notes into EHR predictions that outperform structured data","Clinical notes, not just codes, give EHR models a portability boost","Free-text notes prove key to EHR decision support","Language models read clinical notes better than structured EHR fields","EHR AI portability rises when notes are mined by LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hypothetical numbers in the results tables—such as Clinical ModernBERT at 0.91 AUC against XGBoost's 0.85—are representative of what real experiments on real EHR data would find.","fun_headline_variants_meta":{"raw":{"variants":["LLMs turn clinical notes into EHR predictions that outperform structured data","Clinical notes, not just codes, give EHR models a portability boost","Free-text notes prove key to EHR decision support","Language models read clinical notes better than structured EHR fields","EHR AI portability rises when notes are mined by LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000685,"raw_usage":{"total_tokens":3049,"prompt_tokens":828,"completion_tokens":2221,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":2137}},"tokens_in":444,"tokens_out":2221,"duration_ms":15122,"temperature":1.0,"reasoning_tokens":2137,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:56:03.007596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's claimed comparison on a real multi-institution EHR corpus: train Clinical ModernBERT on text plus ICD codes from one hospital and XGBoost on structured features from the same hospital, then test both on a held-out hospital's data. The central claim fails if the text-based model does not beat the structured-feature baseline on the external site by a clinically meaningful margin, or if the text embedding shows the same institutional shift as the code distributions.","supporting_citations":[],"review_version":1}