Pith. sign in

REVIEW 3 major objections 5 minor 38 references

Bidirectional Representation Learning from Transformers using Multimodal Electronic Health Record Data to Predict Depression

T0 review · 3 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read This paper claims that a bidirectional transformer pretrained on five EHR data types predicts depression better than prior baselines, lifting PRAUC from 0.70 to 0.76.

desk verdict A competent BEHRT-style extension to five EHR modalities, but the depression label is built from the same medication and note features the model sees, so the reported PRAUC gains are not evidence of predictive ability. read the letter →

arxiv 2009.12656 v4 pith:26APHTUP submitted 2020-09-26 cs.LG

classification cs.LG
keywords electronichealthrecordsdepressionpredictiontransformerbidirectionalrepresentationlearningpretrainingandfine-tuningmultimodaldataself-attentioninterpretabilitymaskedlanguagemodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a transformer model trained with masked-code pretraining and then fine-tuned can predict a future depression diagnosis from multimodal electronic health record data better than earlier sequence models. It combines diagnosis codes, procedure codes, medications, demographics, and clinical-note topics into one temporal sequence. On an internal cohort, the model raises precision-recall area under the curve (PRAUC) from 0.70 to 0.76 compared with the best prior baseline, with statistically significant gains at two-week, three-month, six-month, and one-year prediction windows. The authors also claim that self-attention weights expose clinically sensible associations between codes, giving the model an interpretability channel.

What carries the argument

The load-bearing mechanism is the masked-language-model pretraining stage over flattened EHR code sequences, followed by fine-tuning with a classification head. Each code receives token, position, segment, age, and gender embeddings summed into one vector, and bidirectional transformer layers then learn contextual associations. The pretraining objective forces the model to predict randomly masked codes from both left and right context, and the resulting attention weights double as a quantitative map of code-code relations. This carries the claimed performance gain over forward-only and non-pretrained baselines.

What would settle it

Retrain BRLTM on a cohort where depression is confirmed by a later PHQ-9 score or a diagnosis code from a future visit, and remove from the input any medication and note-topic features that mention antidepressants; if PRAUC falls back to the HCET baseline, the reported improvement is largely label leakage rather than genuine temporal prediction.

Watch

Extended reading notes

Core claim

The central claim is that bidirectional representation learning across five EHR modalities improves depression prediction, and that the same learned representations make code-level associations visible. The model, named BRLTM (Bidirectional Representation Learning from Transformers using Multimodal EHR), flattens patient visits into a sequence of codes, adds five kinds of embeddings (code, position, visit segment, age, gender), pretrains on a masked-code prediction task, and then fine-tunes a classification head for depression. Across all four prediction windows and across the three primary diagnoses studied, BRLTM reports the highest ROCAUC and PRAUC, including the headline gain of PRAUC from 0.70 to 0.76 over the best baseline. The paper treats this as evidence that the transfer-learning paradigm and bidirectional attention, previously successful in language modeling, transfer to EHR time series.

Load-bearing premise

The depression label is defined by whether the record already contains depression codes, antidepressants, or antidepressant mentions, and the model's inputs include exactly those medications and note-derived topics, so the label and the features are not independent; if that overlap drives the predictions, the reported gains are not forecasting future disease.

Editorial extensions

If this is right

  • A pretrained multimodal EHR transformer can be fine-tuned on smaller institution-specific datasets, allowing institutions to share a pretrained feature extractor without sharing raw patient records.
  • Including clinical-note topics and demographics, not just diagnosis and procedure codes, contributes to the prediction gain, so future EHR models should keep text-derived features in the input.
  • Attention weights give clinicians a concrete, code-level view of why a patient was flagged, moving closer to usable decision support than black-box prediction.
  • Prediction accuracy falls as the forecast horizon grows from two weeks to one year, meaning the model's practical screening range is short-to-medium term rather than long-term risk forecasting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the depression label includes antidepressants in the medication list or notes, and the model's inputs include exactly those medication codes and note-derived topics, part of the reported PRAUC gain may be the model recognizing the label's own signal; a prospective validation with an independent depression label would separate forecasting from leakage.
  • The same architecture likely transfers to other chronic diseases with similarly patterned temporal EHR signals, but the paper only tests depression, so cross-disease performance remains an open question.
  • Replacing the topic-model features with contextual representations of clinical notes could either raise accuracy further or make the label-overlap problem more severe, depending on how the depression label is defined.
  • Attention maps are suggestive associations rather than causal evidence; clinical deployment would need additional validation before using them as explanations of individual risk.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes BRLTM, a transformer-based bidirectional representation learning model for electronic health records (EHRs), integrating five modalities: diagnoses, procedures, medications, demographics, and topic-model features from clinical notes. The model is pretrained with masked language modeling on EHR sequences and then fine-tuned to predict future diagnosis of depression at four prediction windows. The authors report state-of-the-art performance, with PRAUC increasing from 0.70 (Dipole) or 0.73 (HCET) to 0.78 at the two-week window, and claim statistically significant improvements over the best baseline, HCET. They also present self-attention visualizations as evidence of interpretability and claim that bidirectional learning outperforms forward-only learning.

Significance. If the results were valid, the paper would make a useful contribution by extending BERT-style pretraining and fine-tuning to multimodal, temporally structured EHR data and by demonstrating the value of combining structured codes with clinical-note topic features. Strengths include the release of source code, use of multiple prediction windows, statistically paired comparisons, and attention-based interpretability analysis. However, the central evaluation is undermined by a label-construction scheme that makes the outcome a function of the input features, so the reported PRAUC gains cannot be interpreted as evidence of predictive ability for future depression diagnosis. The significance of the work therefore hinges on whether the leakage can be credibly addressed, which the current manuscript does not do.

major comments (3)
  1. [Section III and Section IV-A, Eq. (2)] The depression label is defined by three criteria: depression-related ICD-9 codes, inclusion of an antidepressant drug in the medication list, or appearance of an antidepressant drug in clinical notes. The model inputs include medication codes (M) and topic features (T) derived from clinical notes, as shown in Eq. (2). For any patient labeled positive because an antidepressant appears in the medication list or a note mentions an antidepressant, that medication code or note-derived topic is present in the input sequence whenever it occurred before the 15-day-excluded window. The 15-day exclusion in Section III removes only the immediate pre-diagnosis period; it does not remove earlier occurrences of the same label-defining codes. Therefore the model can learn a trivial rule detecting its own label, and the PRAUC improvements reported in Table V (e.g., 0.78 vs. 0.73 at two weeks) may reflect shortcut learning rather than forward prediction. This target leakage invalidates the central claim that BRLTM predicts future depression.
  2. [Section V-B, Section IV-E, and Section VII] The conclusion that bidirectional learning is superior to forward-only learning is not supported by a controlled comparison. The best baseline HCET is a different architecture (hierarchical clinical embedding) rather than a forward-only version of BRLTM with identical inputs and training, and the bidirectional baseline Dipole uses only ICD-9 and CPT codes, omitting medications and topics. No ablation fixes the architecture and input modalities while varying only the directionality of the sequence model. Consequently, the statement in Section VII that 'bidirectional learning can provide superior performance to unidirectional' is not demonstrated by the experiments.
  3. [Section IV-B and Table IV] The ablation study for data modality contribution reports only masked-language-modeling precision for different data combinations, not the downstream depression-prediction performance. Table IV shows MLM precision ranging from 0.4248 to 0.5086 as modalities are removed, but no PRAUC or ROCAUC values for the fine-tuned classification task are given. Therefore the Discussion's claim that the results 'emphasize the critical contribution of clinical notes' is not backed by any finetuned-task evidence; the MLM precision numbers measure the pretraining objective, not predictive utility.
minor comments (5)
  1. [Throughout] The manuscript contains numerous typographical errors (e.g., 'sequnetial', 'attetion', 'insitutions', 'aproach') that should be corrected in a revision.
  2. [Section IV-E and Table V] The model name 'BERHT' is spelled inconsistently; the cited work is BEHRT [23]. Please use a single consistent spelling.
  3. [Table III] The number of breast cancer patients in the pretraining cohort is listed as '23,3077', which appears to be a typo; please correct this value.
  4. [Section IV-D] The hyperparameter random search ranges and the LDA topic model settings are not specified; this limits reproducibility despite the availability of source code.
  5. [Fig. 3] The attention visualization examples would be more informative if the EHR code identifiers were explicitly labeled in the figure, since the surrounding text refers to specific codes that are difficult to identify.

Circularity Check

2 steps flagged · score 7.0 of 10

Depression label is defined by the same medication, note, and ICD-code features the model uses as inputs; the reported PRAUC gain partly reflects the model reading its own label, not an independent future prediction.

  1. self definitional [Section III (Data Preprocessing) and Section IV-A, Eq. (2)]
    "depression onset was identified by three methods: depression related ICD-9 codes, inclusion of an antidepressant drug in a patient’s medication list, or appearance of an antidepressant drug in clinical notes / 𝑉𝑡: (𝑋1, 𝑋2, … , 𝑋𝑚𝑡), 𝑋 𝜖 {𝐷, 𝐶, 𝑀, 𝑇} / Any data within 15 days prior to the diagnosis time was excluded to ensure the predictive power of the model."

    The positive label is defined by the presence of depression ICD-9 codes (D), an antidepressant in the medication list (M), or an antidepressant mention in clinical notes (the source of the topic features T). Eq. (2) feeds exactly D, M, and T into the model as input codes. The stated exclusion removes only the 15 days immediately before the assigned diagnosis time, not earlier occurrences of the same label-defining codes in the six-month observation windows. A model can therefore learn the rule “antidepressant medication or depression-related code in input ⇒ positive,” which is the label read back from the feature vocabulary. The reported PRAUC improvement (0.70 to 0.76) thus partly measures recognition of label-defining features rather than forecasting an independent future diagnosis.

  2. other [Section IV-B (Pretraining with MLM), Table III, Section IV-D (training details)]
    "The data used for MLM is shown in the second column of Table III / The dataset for Finetuning of the prediction task underwent 10 random data splits: 70% training, 10% validation, and 20% test. Table III lists pretraining cohorts (10,616 MI; 23,307 breast cancer; 11,757 cirrhosis) and finetuning cohorts (2,943; 5,568; 2,218) drawn from the same patients, without stating that test patients were held out of pretraining."

    Pretraining and fine-tuning use overlapping patients, and the MLM objective randomly masks 15% of codes, including the ICD-9 depression codes that define the outcome. The pretrained model is therefore optimized to predict each patient’s label code from that patient’s own EHR sequence, and the subsequent fine-tuning evaluation is not a clean out-of-sample test. The reported gains over baselines can partly reflect the model having seen the outcome during pretraining rather than an independent ability to forecast depression.

full rationale

The central claim is that BRLTM predicts future depression better than baselines (Section V-B, Table V). This requires the outcome label to encode information not already present in the input features. The paper defines depression onset by depression ICD-9 codes, antidepressants in the medication list, or antidepressant mentions in clinical notes, while the model’s input vocabulary (Eq. (2)) contains diagnosis codes D, medication codes M, and topic features T derived from those same notes. The 15-day exclusion before the assigned diagnosis date does not remove earlier occurrences of the same label-defining codes from the six-month input windows, so the predictor can recover the label directly from the input vocabulary. This is a self-definitional overlap, not merely a strong correlation. In addition, the MLM pretraining uses the same cohort (Table III) and can mask the depression ICD codes that define the label, meaning the fine-tuning evaluation is not fully out-of-sample. The self-citation to HCET [4] for the label definition is not itself circular, since the phenotype definition is an input choice; however, it does not resolve the label/feature overlap. No external validation against an independent depression outcome is reported, so the PRAUC gain cannot be attributed to genuine forward prediction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central performance claim rests on a nontrivial label definition that overlaps with input modalities, plus the transferability of BERT pretraining to EHR. Hyperparameters and LDA settings are fitted choices rather than derived constants.

free parameters (2)
  • MLM hyperparameter configuration = embedding size 216-264, attention layers 6-9, intermediate layer 256-512
    Chosen by random search on the pretraining task (Section IV-B), not derived; downstream performance depends on these choices.
  • LDA topic model settings = 100 topics, 1,000 MALLET iterations
    Clinical notes are compressed into 100 topic features using LDA; this is a heuristic modeling choice that defines inputs.
assumptions (4)
  • domain assumption EHR code sequences can be treated as word sequences for BERT-style representation learning.
    Section IV-B transfers masked language modeling from NLP to EHR without independent evidence that this assumption holds.
  • ad hoc to paper The three-criteria depression proxy is a valid outcome label.
    Section III defines depression onset via ICD-9, antidepressant medication, or antidepressant mention; this overlaps with model inputs.
  • domain assumption Masked language modeling pretraining improves downstream depression prediction.
    The two-stage pretraining and finetuning benefit is assumed from NLP results, not proven by an ablation against random initialization.
  • domain assumption LDA topics meaningfully summarize clinical notes.
    Section III uses LDA with bag-of-words; the paper itself notes this limitation in the Discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bidirectional Representation Learning from Transformers using Multimodal Electronic Health Record Data to Predict Depression." pith.science (2026). https://pith.science/paper/26APHTUP

@misc{pith2026200912656,
  author       = {Pith},
  title        = {Pith review of: Bidirectional Representation Learning from Transformers using Multimodal Electronic Health Record Data to Predict Depression},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/26APHTUP}},
  note         = {Machine review of arXiv:2009.12656}
}
read the original abstract

Advancements in machine learning algorithms have had a beneficial impact on representation learning, classification, and prediction models built using electronic health record (EHR) data. Effort has been put both on increasing models' overall performance as well as improving their interpretability, particularly regarding the decision-making process. In this study, we present a temporal deep learning model to perform bidirectional representation learning on EHR sequences with a transformer architecture to predict future diagnosis of depression. This model is able to aggregate five heterogenous and high-dimensional data sources from the EHR and process them in a temporal manner for chronic disease prediction at various prediction windows. We applied the current trend of pretraining and fine-tuning on EHR data to outperform the current state-of-the-art in chronic disease prediction, and to demonstrate the underlying relation between EHR codes in the sequence. The model generated the highest increases of precision-recall area under the curve (PRAUC) from 0.70 to 0.76 in depression prediction compared to the best baseline model. Furthermore, the self-attention weights in each sequence quantitatively demonstrated the inner relationship between various codes, which improved the model's interpretability. These results demonstrate the model's ability to utilize heterogeneous EHR data to predict depression while achieving high accuracy and interpretability, which may facilitate constructing clinical decision support systems in the future for chronic disease screening and early detection.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 38 canonical work pages

  1. [1]

    Birkhead, Michael Klompas, and Nirav R

    G. S. Birkhead, M. Klompas, and N. R. Shah, “Uses of electronic health records for public health surveillance to advance public health,” Annu. Rev. Public Health, vol. 36, pp. 345–359, 2015, doi: 10.1146/annurev-publhealth-031914-122747

  2. [2]

    Adoption of Electronic Health Record Systems among U.S. Non-Federal Acute Care Hospitals: 2008-2015,

    S. T. & P. V. Henry, J., Pylypchuk, Y., “Adoption of Electronic Health Record Systems among U.S. Non-Federal Acute Care Hospitals: 2008-2015,” ONC Data Brief, no.35., no. 35, pp. 2008– 2015, 2016

  3. [3]

    A new asymp- totic analysis technique for diversity receptions over cor related lognor- mal fading channels

    B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, “Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis,” IEEE J. Biomed. Heal. Informatics, vol. 22, no. 5, pp. 1589–1604, 2018, doi: 10.1109/JBHI.2017.2767063

  4. [4]

    M; Kim, S

    Y. Meng, W. Speier, M. Ong, and C. W. Arnold, “HCET : Hierarchical Clinical Embedding with Topic Modeling on Electronic Health Record for Predicting Depression,” IEEE J. Biomed. Heal. Informatics, 2020, doi: 10.1109/JBHI.2020.3004072

  5. [5]

    RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism,

    E. Choi., M. T. Bahadori, J. A. Kulas, A. Schuetz, W. F. Stewart, and J. Sun, “RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism,” Adv. Neural Inf. Process. Syst. 29 (NIPS 2016), 2016, doi: 10.1063/1.859355

  6. [6]

    Chawla, and Ananthram Swami

    F. Ma, R. Chitta, J. Zhou, Q. You, T. Sun, and J. Gao, “Dipole: Diagnosis prediction in healthcare via attention-based bidirectional recurrent neural networks,” 2017, doi: 10.1145/3097983.3098088

  7. [7]

    A Machine Learning Approach to Classifying Self- Reported Health Status in a Cohort of Patients with Heart Disease Using Activity Tracker Data,

    Y. Meng et al., “A Machine Learning Approach to Classifying Self- Reported Health Status in a Cohort of Patients with Heart Disease Using Activity Tracker Data,” IEEE J. Biomed. Heal. Informatics, vol. 24, no. 3, pp. 878–884, 2020, doi: 10.1109/JBHI.2019.2922178

  8. [8]

    S. L. James et al., “Global, regional, and national incidence, prevalence, and years lived with disability for 354 Diseases and Injuries for 195 countries and territories, 1990-2017: A systematic analysis for the Global Burden of Disease Study 2017,” Lancet, vol. 392, no. 10159, pp. 1789–1858, 2018, doi: 10.1016/S0140- 6736(18)32279-7

Show all 38 references
  1. [9]

    National rates and patterns of depression screening in primary care: Results from 2012 and 2013,

    A. Akincigil and E. B. Matthews, “National rates and patterns of depression screening in primary care: Results from 2012 and 2013,” Psychiatr. Serv., vol. 68, no. 7, pp. 660–666, 2017, doi: 10.1176/appi.ps.201600096

  2. [10]

    Depressive Symptoms as Relative and Attributable Risk Factors for First-Onset Major Depression,

    E. Horwath, J. Johnson, G. L. Klerman, and M. M. Weissman, “Depressive Symptoms as Relative and Attributable Risk Factors for First-Onset Major Depression,” Arch. Gen. Psychiatry, vol. 49, no. 10, p. 817, Oct. 1992, doi: 10.1001/archpsyc.1992.01820100061011

  3. [11]

    The Economic Burden of Adults With Major Depressive Disorder in the United States (2005 and 2010),

    P. E. Greenberg, A.-A. Fournier, T. Sisitsky, C. T. Pike, and R. C. Kessler, “The Economic Burden of Adults With Major Depressive Disorder in the United States (2005 and 2010),” J Clin Psychiatry, vol. 76, no. 2, pp. 155–162, 2015, doi: 10.4088/JCP.14m09298

  4. [12]

    Clinical diagnosis of depression in primary care : a meta-analysis,

    A. J. Mitchell, A. Vaze, S. Rao, and R. Infi, “Clinical diagnosis of depression in primary care : a meta-analysis,” Lancet, vol. 374, no. 9690, pp. 609–619, 2009, doi: 10.1016/S0140-6736(09)60879-5

  5. [13]

    Toward personalizing treatment for depression: predicting diagnosis and severity,

    S. H. Huang, P. LePendu, S. V Iyer, M. Tai-Seale, D. Carrell, and N. H. Shah, “Toward personalizing treatment for depression: predicting diagnosis and severity,” J. Am. Med. Informatics Assoc., vol. 21, no. 6, pp. 1069–1075, Nov. 2014, doi: 10.1136/amiajnl- 2014-002733

  6. [14]

    Development of a Clinical Forecasting Model to Predict Comorbid Depression Among Diabetes Patients and an Application in Depression Screening Policy Making,

    H. Jin, S. Wu, and P. Di Capua, “Development of a Clinical Forecasting Model to Predict Comorbid Depression Among Diabetes Patients and an Application in Depression Screening Policy Making,” Prev. Chronic Dis., vol. 12, pp. 1–10, 2015, doi: 10.5888/pcd12.150047

  7. [15]

    M-SEQ: Early detection of anxiety and depression via temporal orders of diagnoses in electronic health data,

    J. Zhang, H. Xiong, Y. Huang, H. Wu, K. Leach, and L. E. Barnes, “M-SEQ: Early detection of anxiety and depression via temporal orders of diagnoses in electronic health data,” Proc. - 2015 IEEE Int. Conf. Big Data, IEEE Big Data 2015, pp. 2569–2577, 2015, doi: 10.1109/BigData....

  8. [16]

    Mime: Multilevel medical embedding of electronic health records for predictive healthcare,

    E. Choi, C. Xiao, J. Sun, and W. F. Stewart, “Mime: Multilevel medical embedding of electronic health records for predictive healthcare,” in Advances in Neural Information Processing Systems, 2018, pp. 4547–4557

  9. [17]

    Attention Is All You Need,

    A. Vaswani et al., “Attention Is All You Need,” Adv. Neural Inf. Process. Syst., no. Nips, pp. 5998–6008, 2017

  10. [18]

    Learning to Diagnose with LSTM Recurrent Neural Networks,

    Z. C. Lipton, D. C. Kale, C. Elkan, and R. Wetzel, “Learning to Diagnose with LSTM Recurrent Neural Networks,” pp. 1–18, 2015, doi: 10.14722/ndss.2015.23268

  11. [19]

    Doctor AI: Predicting Clinical Events via Recurrent Neural Networks.,

    E. Choi, M. T. Bahadori, A. Schuetz, W. F. Stewart, and J. Sun, “Doctor AI: Predicting Clinical Events via Recurrent Neural Networks.,” Proc. Mach. Learn. Healthc. 2016, vol. 56, pp. 301– 318, 2016

  12. [20]

    Interpretable representation learning for healthcare via capturing disease progression through time,

    T. Bai, B. L. Egleston, S. Zhang, and S. Vucetic, “Interpretable representation learning for healthcare via capturing disease progression through time,” Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 43–51, 2018, doi: 10.1145/3219819.3219904

  13. [21]

    Clinical Intervention Prediction and Understanding using Deep Networks,

    H. Suresh, N. Hunt, A. Johnson, L. A. Celi, P. Szolovits, and M. Ghassemi, “Clinical Intervention Prediction and Understanding using Deep Networks,” arXiv Prepr. arXiv1705.08498v1, pp. 1–16, 2017

  14. [22]

    Learning the Graphical Structure of Electronic Health Records with Graph Convolutional Transformer,

    E. Choi et al., “Learning the Graphical Structure of Electronic Health Records with Graph Convolutional Transformer,” Proc. AAAI Conf. Artif. Intell., vol. 34, no. 01, pp. 606–613, 2020, doi: 10.1609/aaai.v34i01.5400

  15. [23]

    BEHRT: Transformer for Electronic Health Records,

    Y. Li et al., “BEHRT: Transformer for Electronic Health Records,” Sci. Rep., vol. 10, no. 1, pp. 1–17, 2020, doi: 10.1038/s41598-020- 62922-y

  16. [24]

    BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding,” arXiv:11810.04805, 2018

  17. [25]

    Combining billing codes, clinical notes, and medications from electronic health records provides superior phenotyping performance,

    W. Q. Wei, P. L. Teixeira, H. Mo, R. M. Cronin, J. L. Warner, and J. C. Denny, “Combining billing codes, clinical notes, and medications from electronic health records provides superior phenotyping performance,” J. Am. Med. Informatics Assoc., vol. 23, no. e1, pp. 20–27, 2016,...

  18. [26]

    A topic model of clinical reports,

    C. Arnold and W. Speier, “A topic model of clinical reports,” in Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval, 2012, pp. 1031– 1032

  19. [27]

    Evaluating topic model interpretability from a primary care physician perspective,

    C. W. Arnold, A. Oh, S. Chen, and W. Speier, “Evaluating topic model interpretability from a primary care physician perspective,” Comput. Methods Programs Biomed., vol. 124, pp. 67–75, 2015

  20. [28]

    Using phrases and document metadata to improve topic modeling of clinical reports,

    W. Speier, M. K. M. K. Ong, and C. W. C. W. Arnold, “Using phrases and document metadata to improve topic modeling of clinical reports,” J. Biomed. Inform., vol. 61, pp. 260–266, 2016, doi: 10.1016/j.jbi.2016.04.005

  21. [29]

    pandas: a Foundational Python Library for Data Analysis and Statistics,

    W. McKinney, “pandas: a Foundational Python Library for Data Analysis and Statistics,” Python High Perform. Sci. Comput., vol. 14, no. 9, pp. 583–591, 2011, doi: 10.1002/mmce.20381

  22. [30]

    {MALLET: A Machine Learning for Language Toolkit},

    A. K. McCallum, “{MALLET: A Machine Learning for Language Toolkit},” 2002

  23. [31]

    The PHQ-9 : A New Depression Measure,

    M. Kurt Kroenke, MD; Robert L. Spitzer, “The PHQ-9 : A New Depression Measure,” Psychiatr. Ann., vol. 32, no. 9, pp. 509–515, 2002

  24. [32]

    ImageNet Large Scale Visual Recognition Challenge,

    O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, 2015, doi: 10.1007/s11263-015-0816-y

  25. [33]

    Language Models are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog, vol. 1, no. 8, 2019

  26. [34]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015, 2015, pp. 1–15

  27. [35]

    Diabetes mellitus and arthritis: Is it a risk factor or comorbidity?,

    Q. Dong, H. Liu, D. Yang, and Y. Zhang, “Diabetes mellitus and arthritis: Is it a risk factor or comorbidity?,” Med. (United States), vol. 96, no. 18, pp. 1–6, 2017, doi: 10.1097/MD.0000000000006627

  28. [36]

    On Calibration of Modern Neural Networks,

    C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in 34th International Conference on Machine Learning, 2017, pp. 1321–1330

  29. [37]

    Evaluation of Dataset Selection for Pre- Training and Fine-Tuning Transformer Language Models for Clinical Question Answering,

    S. Soni and K. Roberts, “Evaluation of Dataset Selection for Pre- Training and Fine-Tuning Transformer Language Models for Clinical Question Answering,” in 12th Conference on Language Resources and Evaluation (LREC 2020), 2020, no. May, pp. 5532– 5538

  30. [38]

    Federated pretraining and fine tuning of BERT using clinical notes from multiple silos,

    D. Liu and T. Miller, “Federated pretraining and fine tuning of BERT using clinical notes from multiple silos,” arXiv:2002.08562, 2020

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.