Pith. sign in

REVIEW 4 major objections 4 minor 98 references

BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment

T0 review · 4 major / 4 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Aligning LMs with UMLS subgraphs boosts biomedical QA and linking.

desk verdict BALI is a solid pre-training method for biomedical LMs, with real QA gains; the entity-linking improvements are partly circular and need an MLM-only control. read the letter →

arxiv 2509.07588 v1 pith:SAMYO6C5 submitted 2025-09-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords biomedicalknowledgegraphUMLScontrastivelearninglanguagemodelpre-trainingentitylinkingquestionansweringneuralnetworkInfoNCE
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a single additional pre-training stage can inject biomedical factual knowledge into language models by aligning their entity representations with knowledge-graph subgraph representations. Concretely, BALI trains a graph encoder over local UMLS subgraphs and pulls each pooled textual entity mention toward its concept's graph embedding using an InfoNCE contrastive loss, while continuing masked-language modeling. The authors claim that this stage, using only 1.67 million PubMed sentences and about 600,000 UMLS concepts, improves PubMedBERT and BioLinkBERT on question answering and entity linking, and that the graph encoder can then be discarded with no knowledge-graph access needed at inference. If true, this offers a cheap, task-agnostic way to add structured biomedical knowledge to existing encoders.

What carries the argument

Cross-modal anchoring: a single biomedical concept is represented in two complementary modalities — a textual representation (mean-pooled token embeddings of its mention in context) and a structural representation (a GAT-encoded local UMLS subgraph, with initial node features obtained by encoding randomly sampled concept names with the LM). The InfoNCE contrastive loss over paired text-graph anchors is the alignment mechanism, MLM is retained to preserve language ability, and the graph encoder is discarded after pre-training.

What would settle it

Pre-train BALI twice with identical data and compute, once with true mention-to-UMLS links and once with randomly shuffled links; if the shuffled model keeps the QA and entity-linking gains, then the knowledge-graph alignment itself is not the driver and the effect comes from the extra pre-training exposure or MLM.

Watch

Extended reading notes

Core claim

BALI's central claim is that explicit cross-modal alignment between a language model's entity representations and a graph neural network's concept representations transfers usable biomedical knowledge into the LM. For each mention in a masked sentence, the mention's pooled token embeddings serve as the textual anchor, and a GAT encoder produces a structural anchor from a 1-hop UMLS subgraph seeded with concept-name embeddings. An InfoNCE loss pulls paired anchors together while MLM preserves language ability. After this short pre-training, the graph encoder is discarded, yet the LM alone shows mean accuracy gains of 2.1, 1.7, and 6.2 points for PubMedBERT on PubMedQA, MedQA, and BioASQ, larg

Load-bearing premise

The alignment signal is only as good as the automatic linking that matches mentions in the training sentences to UMLS concepts; if many links are wrong, the contrastive objective pushes text embeddings toward the wrong graph nodes, and the reported gains may come mostly from continued masked-language training instead.

Editorial extensions

If this is right

  • Knowledge can be injected during pre-training alone, so downstream tasks need no retrieved subgraphs or entity linking at inference time.
  • A small, balanced alignment corpus (1.67M sentences, ~600K concepts, 65K steps, roughly 9 GPU-hours for base models) is sufficient for consistent QA improvements.
  • Zero-shot entity linking improves sharply for general biomedical LMs, with Accuracy@1 rising by about 13 points for PubMedBERT and 24 points for BioLinkBERT-base on average across five corpora.
  • A BALI-adapted BioLinkBERT-base can match or slightly beat SapBERT, a task-specific model pre-trained on the full UMLS synonym space, on supervised entity linking.
  • GAT-based subgraph aggregation outperforms mean-pooling graph encoders, while larger LMs can benefit more from a single-encoder linearized-graph variant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own ablation shows that removing the alignment loss but keeping MLM already yields 63.78 on PubMedQA versus the 63.1 raw PubMedBERT baseline, so part of the gain is plausibly from continued masked-language pre-training rather than graph alignment; no fully matched equal-compute control is reported.
  • Because the training data relies on automatic entity recognition and normalization, incorrect mention-to-concept links would push text embeddings toward wrong graph nodes; a gold-annotated or confidence-filtered training set would likely sharpen the measured alignment signal.
  • The same anchor-based recipe should transfer to other text-attributed knowledge graphs, with the effective graph encoder choice depending on model capacity — small encoders appear to need the external GAT, larger ones can use linearized subgraphs.
  • A direct negative-control experiment, pairing sentences with randomly shuffled subgraphs instead of their true linked subgraphs, would isolate whether the contrastive alignment itself or the extra pre-training data drives the reported gains.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes BALI, a pre-training method that augments biomedical language models (PubMedBERT, BioLinkBERT base/large) by aligning pooled entity-mention representations with UMLS knowledge-graph subgraph representations. A GAT encoder encodes local 1-hop subgraphs, whose initial node features are LM-encoded concept names; an InfoNCE contrastive loss pulls mention and graph representations of the same concept together, while MLM is retained as a joint objective. After pre-training on 1.67M PubMed sentences with BERN2-based entity linking, the GNN is discarded, and the resulting LM is evaluated on biomedical question answering (PubMedQA, MedQA, BioASQ), entity linking (zero-shot and supervised), and relation extraction (ChemProt, DDI, GAD). The paper reports consistent QA gains over base models, large zero-shot entity-linking gains for general biomedical LMs, small relation-extraction gains, and ablations showing that both objectives and attention-based graph aggregation matter.

Significance. If the claims hold, BALI would be a valuable and relatively cheap recipe for infusing structured biomedical knowledge into lightweight encoders without any inference-time KG access. The paper has concrete strengths: it releases code and pre-trained models, reports hyperparameters in Table 1, averages repeated fine-tuning runs for PubMedQA and BioASQ, and includes a useful ablation suite (Table 6) covering loss removal, alternative alignment losses, aggregation methods, and GNN depth. The core idea of explicit cross-modal alignment through entity anchors is clearly presented and is distinct from prior interaction-token or retrieval-augmented approaches. However, the current evidence for the entity-representation claim is weakened by the absence of an MLM-only control in the entity-linking experiments and by the close similarity between the zero-shot entity-linking protocol and the training objective. The QA results are more independent and, with the existing MLM-only ablation, already show that the alignment term contributes beyond continued pre-training, which is a useful partial control.

major comments (4)
  1. [§4.2, §5.3.2, Table 3] The zero-shot entity-linking evaluation largely re-measures the training objective. L_align (Eq. 5) maximizes cosine similarity between a pooled mention representation and a GNN subgraph representation whose initial node features are LM-encoded concept names (§4.1). The zero-shot retrieval protocol (§5.3.2) retrieves by cosine similarity between mention and concept-name representations. This is close to the positive-pair construction in training. More importantly, Table 3 reports no MLM-only control for entity linking, even though Table 6 shows that MLM-only accounts for a substantial portion of the QA gains (PubMedQA 63.1→63.78; BioASQ 67.8→70.58). Without an equivalent EL control or a disjoint-concept test, the headline EL gains (e.g., PubMedBERT +13.1 Accuracy@1) cannot be cleanly attributed to the KG-alignment term rather than to continued MLM on the same entity-linked corpus.
  2. [§5.3.1, Table 2] The text claims that BioLinkBERT_large 'after BALI pretraining performs on par or better than the task-specific QA-GNN and GreaseLM methods.' This is contradicted by Table 2: BioLinkBERT_large+BALI(GNN) scores 68.7 on PubMedQA and 45.0 on MedQA, while QA-GNN scores 72.1/45.0 and GreaseLM 72.4/45.1. Even the linear-graph variant (70.9 PubMedQA) is below QA-GNN and GreaseLM. The comparison should be corrected, or the claim narrowed to specific datasets where it actually holds (e.g., MedQA).
  3. [§5.1 / Table 2] MedQA results are reported as single means with no standard deviations or error bars, whereas PubMedQA and BioASQ have ± values averaged over 10 runs. Since the reported MedQA gains are small (e.g., PubMedBERT 38.1→39.8; BioLinkBERT_large 44.6→45.0), it is impossible to assess whether these differences are significant. Please report variance or at least the number of seeds used for MedQA.
  4. [§5, 'Pretraining Data'] The pre-training corpus is built with BERN2 for named entity recognition and UMLS normalization, but the paper reports no quality check on this linking. If a substantial fraction of mentions are mapped to incorrect UMLS concepts, the contrastive loss will pull text representations toward wrong graph nodes, and the observed gains would reflect a different effect. The manuscript should at least report a sample-based precision estimate for the BERN2 linking on the pre-training data, or provide a robustness analysis (e.g., training on a filtered subset with high-confidence links).
minor comments (4)
  1. [§5.1] The evaluation-task list has a duplicate: item (iii) is listed as 'BC5CDR-D' but should presumably be the BC5CDR-Chemical corpus, matching Table 3's 'BC5CDR-C' column.
  2. [§5.5.3 / Table 6] The ablation text says 'larger (5 layers)' but Table 6 reports L=7 for the larger GNN. The text and table should be harmonized.
  3. [Abstract / §1 / §5 / Conclusion] The pre-training corpus size is given inconsistently as '1.5M sentences' (Introduction), '1.67M sentences' (Pretraining Data), and '1.7M sentences' (Conclusion). Please use one consistent figure.
  4. [Table 3] The underline convention ('best of two scores') is defined only for original vs. BALI-pretrained models. For the SapBERT and GEBERT rows, the notational significance of underlining is less clear; please clarify or add bolding for the overall best per column.

Circularity Check

1 steps flagged · score 4.0 of 10

Zero-shot entity-linking evaluation largely re-measures the BALI alignment objective; QA/RE provide independent support, so circularity is partial.

  1. fitted input called prediction [Section 4.2 (L_align) and Section 5.3.2 (zero-shot retrieval protocol)]
    "L_align = − 1/B Σ_i log exp(cos(ē_i, ḡ_i)/τ) / Σ_j exp(cos(ē_j, ḡ_j)/τ) ... As an initial representation for a node u, a random concept name s_u ∈ S_u is sampled and encoded with a textual encoder: ḡ_u^(0) = LM(s_u). ... zero-shot similarity-based retrieval approach over pooled mention and concept name representations [74]."

    The zero-shot EL protocol retrieves concepts by cosine similarity between a pooled mention representation and a concept-name representation. The BALI alignment loss L_align directly maximizes cosine similarity between the same pooled mention representation ē_v and a graph representation ḡ_v whose initial node feature is the LM-encoded concept name (ḡ_u^(0)=LM(s_u)). After training, the GNN output is a function of these concept-name embeddings, so the mention representations have been pulled toward concept-name-derived vectors by the training objective itself. Reporting large zero-shot EL gains as evidence that 'external KG structure was internalized' is therefore partly circular: the evaluation metric is essentially the training similarity, not an independent probe of graph structure. The

full rationale

The paper's central claims are supported by a mix of independent and partly circular evidence. The QA evaluations (PubMedQA, MedQA, BioASQ) and relation extraction (ChemProt, DDI, GAD) are external downstream tasks not defined by the alignment objective; the ablation in Table 6 shows that removing L_align (MLM-only) yields 63.78 on PubMedQA and 70.58 on BioASQ vs 63.1/67.8 for the base model, so continued LM pre-training accounts for part of the gain, but the full BALI model (65.2/74) still exceeds the MLM-only control, giving the alignment term some independent content. The zero-shot entity linking results in Table 3 are the main evidence for 'quality of entity representations,' and these are close to the training objective: the model is trained with an InfoNCE loss on cos(mention, graph-concept) and tested with cosine retrieval against concept names, where the graph representation is initialized from LM-encoded concept names. This is a partial reduction by construction, but not a full one because the GNN aggregates neighboring concept-name embeddings and the evaluation corpora are not identical to the pre-training sentences. The absence of an MLM-only control for entity linking is a gap but not itself a circular step. No load-bearing self-citation chain was found: references to the authors' prior GEBERT work ([58]) support GAT choice and MS-loss hyperparameters but are not the basis for the main empirical claim. No uniqueness theorem or ansatz smuggled via citation. Overall, circularity is partial and limited to the entity-linking evaluation; the paper is otherwise self-contained against external benchmarks.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

No new entities are introduced; the method operates on existing UMLS concepts. The main free choices are standard hyperparameters, but the unstated InfoNCE temperature is a notable gap. The key assumptions concern the quality of the pre-training corpus and the transfer of alignment to general LM representations.

free parameters (4)
  • InfoNCE temperature tau = not stated
    The contrastive temperature is not reported; it materially controls the sharpness of the alignment loss.
  • GNN hidden size, layers, heads = 768, 5 layers, 2 heads
    Graph encoder capacity choices, evaluated only partially through layer-count ablation.
  • Max neighbors per node = 3
    Subgraph truncation chosen to limit memory; directly shapes the graph context.
  • LM and non-LM learning rates = 2e-5 and 1e-4
    Standard pre-training hyperparameters; not fitted to downstream tasks but they affect results.
assumptions (3)
  • domain assumption UMLS concept names and relations are a reliable source of biomedical knowledge.
    The entire method assumes aligning text to UMLS subgraphs transfers useful factual knowledge.
  • domain assumption BERN2 entity recognition and normalization on PubMed abstracts is accurate enough for the contrastive pairs.
    Noisy links would corrupt the alignment signal; the paper does not analyze this error rate.
  • ad hoc to paper Aligning pooled entity representations improves the LM more broadly, not just the entity pool.
    This is the core mechanism that the paper assumes but does not directly prove.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment." pith.science (2026). https://pith.science/paper/SAMYO6C5

@misc{pith2026250907588,
  author       = {Pith},
  title        = {Pith review of: BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SAMYO6C5}},
  note         = {Machine review of arXiv:2509.07588}
}
read the original abstract

In recent years, there has been substantial progress in using pretrained Language Models (LMs) on a range of tasks aimed at improving the understanding of biomedical texts. Nonetheless, existing biomedical LLMs show limited comprehension of complex, domain-specific concept structures and the factual information encoded in biomedical Knowledge Graphs (KGs). In this work, we propose BALI (Biomedical Knowledge Graph and Language Model Alignment), a novel joint LM and KG pre-training method that augments an LM with external knowledge by the simultaneous learning of a dedicated KG encoder and aligning the representations of both the LM and the graph. For a given textual sequence, we link biomedical concept mentions to the Unified Medical Language System (UMLS) KG and utilize local KG subgraphs as cross-modal positive samples for these mentions. Our empirical findings indicate that implementing our method on several leading biomedical LMs, such as PubMedBERT and BioLinkBERT, improves their performance on a range of language understanding tasks and the quality of entity representations, even with minimal pre-training on a small alignment dataset sourced from PubMed scientific abstracts.

Figures

Figures reproduced from arXiv: 2509.07588 by the authors.

Figure 1
Figure 1. The overall framework. We first retrieve subgraphs from the knowledge graph based on the entities in a text fragment [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

98 extracted references · 42 canonical work pages

  1. [1]

    Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Work- shop, Anna Rumshisky, Kirk Roberts, Steven Bethard, and Tristan Naumann (Eds.). Association for Computational Linguistics, Minne...

  2. [2]

    Jinheon Baek, Alham Fikri Aji, and Amir Saffari. 2023. Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answer- ing. In Proceedings of the 1st Workshop on Natural Language Reasoning and Struc- tured Explanations (NLRSE). 78–106

  3. [3]

    Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, and Yulia Tsvetkov. 2024. Kgquiz: Evaluating the generalization of encoded knowledge in large language models. In Proceedings of the ACM on Web Conference 2024. 2226–2237

  4. [4]

    Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text. InProceedings of the 2019 Conference on Empirical Meth- ods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association...

  5. [5]

    Olivier Bodenreider. 2004. The Unified Medical Language System (UMLS): inte- grating biomedical terminology. Nucleic Acids Research 32, Database-Issue (2004), 267–270. doi:10.1093/NAR/GKH061

  6. [6]

    Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, Un...

  7. [7]

    Àlex Bravo, Janet Piñero, Núria Queralt-Rosinach, Michael Rautschka, and Laura I Furlong. 2015. Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research. BMC bioinfor- matics 16 (2015), 1–17

  8. [8]

    Shaked Brody, Uri Alon, and Eran Yahav. 2022. How Attentive are Graph At- tention Networks?. In The Tenth International Conference on Learning Repre- sentations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net. https: //openreview.net/forum?id=F72ximsx7C1

Show all 98 references
  1. [9]

    David Chang, Ivana Balazevic, Carl Allen, Daniel Chawla, Cynthia Brandt, and Richard Andrew Taylor. 2020. Benchmark and Best Practices for Biomedical Knowledge Graph Embeddings. InProceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, BioNLP 2020, Online...

  2. [10]

    Qijie Chen, Haotong Sun, Haoyang Liu, Yinghui Jiang, Ting Ran, Xurui Jin, Xianglu Xiao, Zhimin Lin, Hongming Chen, and Zhangmin Niu. 2023. An extensive benchmark study on biomedical text generation and min- ing with ChatGPT. Bioinformatics 39, 9 (09 2023), btad557. doi:10.1093...

  3. [11]

    David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Ko- zlova, Nikita Semenov, and Alexander Panchenko. 2021. Text Detoxification using Large Pre-trained Neural Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing ...

  4. [12]

    Allan Peter Davis, Thomas C Wiegers, Michael C Rosenstein, and Carolyn J Mattingly. 2012. MEDIC: a practical disease vocabulary used at the Comparative Toxicogenomics Database. Database 2012 (2012), bar065

  5. [13]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Hum...

  6. [14]

    Rezarta Islamaj Dogan, Robert Leaman, and Zhiyong Lu. 2014. NCBI disease corpus: A resource for disease name recognition and concept normalization. J. Biomed. Informatics 47 (2014), 1–10. doi:10.1016/J.JBI.2013.12.006

  7. [15]

    Matthias Fey and Jan E. Lenssen. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds

  8. [16]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural Message Passing for Quantum Chemistry. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedin...

  9. [17]

    Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2022. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans. Comput. Heal. 3, 1 (2022), 2:1–2:23. doi:10.1145/3458754

  10. [18]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning. PMLR, 3929–3938

  11. [19]

    Hamilton, Zhitao Ying, and Jure Leskovec

    William L. Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Represen- tation Learning on Large Graphs. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , I...

  12. [20]

    Bin He, Di Zhou, Jinghui Xiao, Xin Jiang, Qun Liu, Nicholas Jing Yuan, and Tong Xu. 2020. Integrating Graph Contextualized Knowledge into Pre-trained Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (F...

  13. [21]

    María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry De- clerck. 2013. The DDI corpus: An annotated corpus with pharmacological sub- stances and drug–drug interactions. Journal of biomedical informatics 46, 5 (2013), 914–920

  14. [22]

    Chun-Yu Hsueh, Yu Zhang, Yu-Wei Lu, Jen-Chieh Han, Wilailack Meesawad, and Richard Tzong-Han Tsai. 2023. NCU-IISR: Prompt Engineering on GPT-4 to Stove Biological Problems in BioASQ 11b Phase B. In Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), T...

  15. [23]

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences 11, 14 (2021). doi:10.3390/app11146421

  16. [24]

    Cohen, and Xinghua Lu

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu

  17. [25]

    Minki Kang, Jinheon Baek, and Sung Ju Hwang. 2022. KALA: Knowledge- Augmented Language Model Adaptation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies, NAACL 2022, Seattle, W ...

  18. [26]

    Seyed Mehran Kazemi and David Poole. 2018. SimplE Embedding for Link Prediction in Knowledge Graphs. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada , S...

  19. [27]

    Pei Ke, Haozhe Ji, Yu Ran, Xin Cui, Liwei Wang, Linfeng Song, Xiaoyan Zhu, and Minlie Huang. 2021. JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online...

  20. [28]

    Jing Yu Koh, Daniel Fried, and Russ Salakhutdinov. 2023. Generating Im- ages with Multimodal Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December ...

  21. [29]

    Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding Language Models to Images for Multimodal Inputs and Outputs. InInternational Conference BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment SIGIR ’25, July...

  22. [30]

    Martin Krallinger, Obdulia Rabal, Saber A Akhondi, Martın Pérez Pérez, Jesús Santamaría, Gael Pérez Rodríguez, Georgios Tsatsaronis, Ander Intxaurrondo, José Antonio López, Umesh Nandal, et al . 2017. Overview of the BioCreative VI chemical-protein interaction Track. In Procee...

  23. [31]

    Andrey Kutuzov, Mohammad Dorgham, Oleksiy Oliynyk, Chris Biemann, and Alexander Panchenko. 2019. Learning Graph Embeddings from WordNet-based Similarity Measures. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019) , Rada Mihalcea, ...

  24. [32]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  25. [33]

    Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J

    Jiao Li, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J. Mattingly, Thomas C. Wiegers, and Zhiyong Lu. 2016. BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database J. Biol. Databa...

  26. [34]

    Qing Li, Lei Li, and Yu Li. 2024. Developing ChatGPT for biology and medicine: a complete review of biomedical question answering. Biophysics Reports 10, 3 (2024), 152–171. doi:10.52601/bpr.2024.240004 Received: 2024/01/15; Accepted: 2024/02/19; Published: 2024/06/30

  27. [35]

    Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Col- lier. 2021. Self-Alignment Pretraining for Biomedical Entity Representations. In Proceedings of the 2021 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Huma...

  28. [36]

    Fangyu Liu, Ivan Vulic, Anna Korhonen, and Nigel Collier. 2021. Learning Domain-Specialised Representations for Cross-Lingual Biomedical Entity Link- ing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Join...

  29. [37]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Vi- sual Instruction Tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Al...

  30. [38]

    Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020. K-BERT: Enabling Language Representation with Knowledge Graph. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artif...

  31. [39]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)

  32. [40]

    Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Or- leans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum? id=Bkg6RiCqY7

  33. [41]

    Donna Maglott, Jim Ostell, Kim D Pruitt, and Tatiana Tatusova. 2007. Entrez Gene: gene-centered information at NCBI. Nucleic acids research 35, suppl_1 (2007), D26–D31

  34. [42]

    Aidan Mannion, Didier Schwab, and Lorraine Goeuriot. 2023. UMLS-KGI- BERT: Data-Centric Knowledge Integration in Transformers for Biomedical Entity Recognition. In Proceedings of the 5th Clinical Natural Language Processing Workshop, ClinicalNLP@ACL 2023, Toronto, Canada, July...

  35. [43]

    Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, and Nigel Collier. 2022. Rewire-then-Probe: A Contrastive Recipe for Probing Biomed- ical Knowledge of Pre-trained Language Models. InProceedings of the 60th Annual Meeting of the Association for Computati...

  36. [44]

    Chen, and Alexan- der Wong

    George Michalopoulos, Yuanxin Wang, Hussam Kaka, Helen H. Chen, and Alexan- der Wong. 2021. UmlsBERT: Clinical Domain Knowledge Augmentation of Con- textual Embeddings Using the Unified Medical Language System Metathesaurus. In Proceedings of the 2021 Conference of the North A...

  37. [45]

    Fedor Moiseev, Zhe Dong, Enrique Alfonseca, and Martin Jaggi. 2022. SKILL: Structured Knowledge Infusion for Large Language Models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies...

  38. [46]

    Alexander A Morgan, Zhiyong Lu, Xinglong Wang, Aaron M Cohen, Juliane Fluck, Patrick Ruch, Anna Divoli, Katrin Fundel, Robert Leaman, Jörg Hakenberg, et al. 2008. Overview of BioCreative II gene normalization. Genome biology 9 (2008), 1–19

  39. [47]

    Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Salvador Lima- López, Eulàlia Farré-Maduell, Luis Gascó, Martin Krallinger, and Georgios Paliouras. 2023. Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Questi...

  40. [48]

    Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, and Georgios Paliouras. 2023. Overview of BioASQ Tasks 11b and Synergy11 in CLEF2023. In Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), Thessaloniki, Greece, September 18th to 21st, 20...

  41. [49]

    John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. 2008. Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue 6, 2 (2008), 40–53

  42. [50]

    Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. 2023. Capabilities of GPT-4 on Medical Challenge Problems. CoRR abs/2303.13375 (2023). doi:10.48550/ARXIV.2303.13375 arXiv:2303.13375

  43. [51]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023). doi:10. 48550/ARXIV.2303.08774 arXiv:2303.08774

  44. [52]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Des- maison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...

  45. [53]

    Peters, Mark Neumann, Robert L

    Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A. Smith. 2019. Knowledge Enhanced Contextual Word Representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inte...

  46. [54]

    Phan, Aixin Sun, and Yi Tay

    Minh C. Phan, Aixin Sun, and Yi Tay. 2019. Robust Representation Learning of Biomedical Names. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, Anna Korhonen, Davi...

  47. [55]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, Virtual Event...

  48. [56]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deep- Speed: System Optimizations Enable Training Deep Learning Models with Over SIGIR ’25, July 13–18, 2025, Padua, Italy Andrey Sakhovskiy and Elena Tutubalina 100 Billion Parameters. In KDD ’20: The 26t...

  49. [57]

    Bennett, and Saurabh Tiwary

    Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul N. Bennett, and Saurabh Tiwary. 2020. Knowledge-Aware Language Model Pretraining. CoRR abs/2007.00655 (2020). arXiv:2007.00655 https://arxiv.org/abs/2007.00655

  50. [59]

    doi:10.1109/SC41405.2020.00024

  51. [60]

    Mikhail Salnikov, Hai Le, Prateek Rajput, Irina Nikishina, Pavel Braslavski, Valentin Malykh, and Alexander Panchenko. 2023. Large Language Models Meet Knowledge Graphs to Answer Factoid Questions. In Proceedings of the 37th Pacific Asia Conference on Language, Information and...

  52. [61]

    Mohammad, Goran Nenadic, and Graciela Gonzalez-Hernandez

    Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kir- itchenko, Farrokh Mehryary, Sifei Han, Tung Tran, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, and Graciela Gonzalez-Hernandez. 201...

  53. [62]

    Matthias Schildwächter, Alexander Bondarenko, Julian Zenker, Matthias Hagen, Chris Biemann, and Alexander Panchenko. 2019. Answering Comparative Ques- tions: Better than Ten-Blue-Links?. InProceedings of the 2019 Conference on Human Information Interaction and Retrieval, CHIIR...

  54. [63]

    Özge Sevgili, Alexander Panchenko, and Chris Biemann. 2019. Improving Neural Entity Disambiguation with Graph Embeddings. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Work- shop, Fernando Alva-Manchego, Eunsol Choi...

  55. [64]

    Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-Augmented Black-Box Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa...

  56. [65]

    Andrey Sakhovskiy, Natalia Semenova, Artur Kadurin, and Elena Tutubalina

  57. [66]

    Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Ji- axiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137 (2021)

  58. [67]

    Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowl- edge Graph Embedding by Relational Rotation in Complex Space. In 7th Interna- tional Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net. https://op...

  59. [68]

    Mujeen Sung, Hwisang Jeon, Jinhyuk Lee, and Jaewoo Kang. 2020. Biomedical Entity Representations with Synonym Marginalization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Dan Jurafsky, Joyce Cha...

  60. [69]

    Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jin- hyuk Lee, and Jaewoo Kang. 2022. BERN2: an advanced neural biomedical named entity recognition and normalization tool. Bioin- formatics 38, 20 (09 2022), 4837–4839. doi:10.1093/bioinformatics/ btac598 arXiv:https://...

  61. [70]

    Yi, Minji Jeon, Sungdong Kim, and Jae- woo Kang

    Mujeen Sung, Jinhyuk Lee, Sean S. Yi, Minji Jeon, Sungdong Kim, and Jae- woo Kang. 2021. Can Language Models be Biomedical Knowledge Bases?. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Domin...

  62. [71]

    Yijun Tian, Huan Song, Zichen Wang, Haozhu Wang, Ziqing Hu, Fang Wang, Nitesh V Chawla, and Panpan Xu. 2024. Graph neural prompting with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19080–19088

  63. [72]

    Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang. 2020. CoLAKE: Contextualized Language and Knowledge Embedding. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online)...

  64. [73]

    Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. 2016. Complex Embeddings for Simple Link Prediction. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016 (JMLR Work...

  65. [74]

    Elena Tutubalina, Artur Kadurin, and Zulfat Miftahutdinov. 2020. Fair Evaluation in Concept Normalization: a Large-scale Comparative Analysis for BERT-based Models. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain ...

  66. [75]

    Elena Tutubalina, Artur Kadurin, and Zulfat Miftahutdinov. 2020. Fair evaluation in concept normalization: a large-scale comparative analysis for BERT-based mod- els. In Proceedings of the 28th International conference on computational linguistics . 6710–6716

  67. [76]

    Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. ArXiv abs/1807.03748 (2018). https://api. semanticscholar.org/CorpusID:49670925

  68. [77]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: An- nual Conference on Neural Information Processing Systems ...

  69. [78]

    Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph Attention Networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedi...

  70. [79]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation ...

  71. [80]

    Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. KEPLER: A Unified Model for Knowledge Embed- ding and Pre-trained Language Representation. Trans. Assoc. Comput. Linguistics 9 (2021), 176–194. doi:10.1162/TACL_A_00360

  72. [81]

    Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R. Scott

  73. [82]

    Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2015. Em- bedding Entities and Relations for Learning and Inference in Knowledge Bases. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track...

  74. [83]

    Manning, Percy Liang, and Jure Leskovec

    Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christo- pher D. Manning, Percy Liang, and Jure Leskovec. 2022. Deep Bidirec- tional Language-Knowledge Graph Pretraining. In Advances in Neural Infor- mation Processing Systems 35: Annual Conference on Neural Info...

  75. [84]

    Michihiro Yasunaga, Jure Leskovec, and Percy Liang. 2022. LinkBERT: Pretraining Language Models with Document Links. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , ...

  76. [85]

    Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational ...

  77. [86]

    Donghan Yu, Chenguang Zhu, Yiming Yang, and Michael Zeng. 2022. JAKET: Joint Pre-training of Knowledge Graph and Language Understanding. In Thirty- Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty-Fourth Confer- ence on Innovative Applications of Artificial ...

  78. [87]

    Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. KEPLER: A Unified Model for Knowledge Embed- ding and Pre-trained Language Representation. Transactions of the Association for Computational Linguistics 9 (2021), 176–194

  79. [88]

    Manning, and Jure Leskovec

    Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D. Manning, and Jure Leskovec. 2022. GreaseLM: Graph REASoning Enhanced Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 2...

  80. [90]

    In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019

    Multi-Similarity Loss With General Pair Weighting for Deep Metric Learn- ing. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 . Computer Vision Foundation / IEEE, 5022–5030. doi:10.1109/CVPR.2019.00516

  81. [96]

    Zheng Yuan, Zhengyun Zhao, Haixia Sun, Jiao Li, Fei Wang, and Sheng Yu. 2022. CODER: Knowledge-infused cross-lingual medical term embedding for term normalization. J. Biomed. Informatics 126 (2022), 103983. doi:10.1016/J.JBI.2021. 103983

  82. [99]

    In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

    ERNIE: Enhanced Language Representation with Informative Entities. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1441–1451

  83. [100]

    Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu

  84. [101]

    ERNIE: Enhanced Language Representation with Informative Entities. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Ko- rhonen, David R. Traum, and Lluís Màr...

  85. [2019]

    PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Process- ing, EMNLP-IJCNLP 2019, Hong Kong, China, November ...

  86. [2023]

    Graph-Enriched Biomedical Entity Representation Transformer. In Experi- mental IR Meets Multilinguality, Multimodality, and Interaction - 14th International Conference of the CLEF Association, CLEF 2023, Thessaloniki, Greece, September 18-21, 2023, Proceedings (Lecture Notes i...

  87. [2080]

    http://proceedings.mlr.press/v48/trouillon16.html

  88. [3506]

    doi:10.1145/3394486.3406703

  89. [4810]

    doi:10.18653/V1/2022.ACL-LONG.329

  90. [5167]

    doi:10.18653/V1/2022.NAACL-MAIN.379

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.