Pith. sign in

REVIEW 3 major objections 6 minor 81 references

Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper is a systematic literature review that assembles 26 adapter-based methods for injecting knowledge into language models, classifies them by architecture, domain, and task, and shows that adapter-based approaches consistently…

desk verdict Useful map of adapter-based KELMs, but the screening counts don't add up — fix that before trusting the 'systematic' claims. read the letter →

arxiv 2411.16403 v1 pith:LGYE6W65 submitted 2024-11-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords knowledge-enhancedlanguagemodelsadaptermodulesknowledgegraphssystematicliteraturereviewbiomedicalnaturalprocessingparameter-efficientfine-tuninginjectionsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a systematic literature review of adapter-based knowledge-enhanced language models (KELMs): large language models that receive structured knowledge, usually from knowledge graphs, through small trainable adapter modules rather than full fine-tuning. It assembles 26 peer-reviewed papers from a structured search and classifies them by adapter architecture, domain scope, knowledge source, and downstream task. It finds that open-domain 'general knowledge' injection and closed-domain biomedical injection have both been actively explored, that the most common adapter architectures are the original bottleneck design and its simplified single-module variant, and that on shared biomedical benchmarks the MoP and KEBLM frameworks deliver the largest accuracy gains over base models. The contribution is a structured map of a young field, intended as an entry point for researchers.

What carries the argument

The central object is the adapter module itself: a small bottleneck feed-forward block inserted into a frozen transformer layer, adding only a few percent of new parameters and trained on external knowledge while the base model stays fixed. The review's machinery is a categorization scheme that sorts the 26 papers along four axes: adapter type (original bottleneck, single-adapter variant, K-Adapter plug-ins, and custom designs), domain scope (open vs. closed, with biomedical singled out), knowledge source (ConceptNet, DBpedia, UMLS, and others), and downstream task (question answering, named entity recognition, reading comprehension, dialogue, and more). This scheme is what turns a scattered set of papers into a comparable landscape, and it supports the paper's head-to-head biomedical performance table.

What would settle it

A reader could rerun the search with the same inclusion criteria plus additional databases such as arXiv and Scopus and broader query synonyms such as 'parameter-efficient' or 'prompt tuning'; if this expanded search surfaces a substantial body of adapter-based KELM papers absent from the 26, the review's claim to comprehensive coverage would be weakened.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that adapter-based knowledge enhancement is not a niche corner of the KELM literature but a recognizably distinct and rapidly growing approach with a clear methodological structure. The review shows that general-knowledge and domain-specific approaches have been frequently explored, identifies the single-adapter-bottleneck configuration as the predominant adapter type, and documents the biomedical domain as the most active closed-domain application area. Its quantitative comparison on five biomedical benchmarks (HoC, PubMedQA, BioASQ7b, MedNLI, and NCBI) reports that adapter-based KELMs consistently improve over their base models, with the MoP and KEBLM frameworks producing the strongest gains and the CPK and DAKI frameworks trailing. The paper's central claim is that prior surveys missed most of this adapter-based work, so this review fills that gap.

Load-bearing premise

The survey's map of the field is only as complete as its literature search, which restricted inclusion to English peer-reviewed papers in three bibliographic databases matching a fixed query and stopped in January 2024, so relevant work using other terms or venues could be missing.

Editorial extensions

If this is right

  • A new researcher can use the classification to locate the adapter architecture and benchmark most likely to suit a target domain.
  • Knowledge-intensive tasks such as question answering and named entity recognition are the established testbeds, while generative tasks like summarization and open-ended text generation remain largely unexplored.
  • In the biomedical domain, MoP and KEBLM are the recommended frameworks when the goal is maximising accuracy on the five shared benchmarks.
  • The roughly linear yearly increase in publications suggests the field will keep expanding with novel adapter architectures.
  • The balance between open-domain and closed-domain approaches, with biomedical dominant among closed domains, points to legal or financial document understanding as likely next applications.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the performance pattern holds beyond the five benchmarks, adapter-based knowledge injection may become the default low-cost route for domain adaptation in regulated fields, since adapters preserve the base model and reduce catastrophic forgetting.
  • The review's exclusion of LoRA/QLoRA from the core pool, despite noting them, hints that the boundary between 'adapter layers' and 'low-rank weight updates' is becoming blurred; a future survey might treat them under one parameter-efficient umbrella.
  • A testable extension is to run a controlled comparison of the four adapter types on a single set of knowledge-intensive tasks with identical base models and knowledge sources, since the paper's cross-paper comparison cannot fully control for training data and hyperparameters.
  • The biomedical dominance may reflect the availability of UMLS and similar curated knowledge graphs, so applying the same adapter recipe to legal or financial knowledge graphs could transfer the observed gains.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript presents a systematic literature review (SLR) of adapter-based approaches to knowledge-enhanced language models (KELMs). The authors identify a gap in prior surveys (Colon-Hernandez et al., 2021; Wei et al., 2021) regarding adapter-based methods and aim to fill it by reviewing 26 papers from ACM, ACL, IEEE, and other sources. The paper categorizes adapter architectures (Houlsby, Pfeiffer, Bapna and Firat, K-Adapter, and unique variants), domains (open vs. closed, with biomedical as the most frequent closed domain), and downstream tasks. It provides quantitative trends (yearly growth, adapter-type distribution, domain distribution) and a qualitative synthesis of general, linguistic, domain-specific, and biomedical knowledge injection approaches. A particular contribution is the biomedical performance comparison in Table 3, which reports accuracy/F1 on five tasks across three base models enhanced with MoP, KEBLM, DAKI, and CPK. The paper concludes with current trends and future directions. The final pool of 26 papers, their categorizations, and the performance comparison form the empirical basis of the survey.

Significance. If the review is reproducible and the reported comparisons are sound, the paper provides a valuable structured entry point to a nascent and fast-growing subfield. The taxonomy of adapter types and the domain/task categorization are useful organizational contributions, and the biomedical performance comparison, though focused, is more specific than what prior KELM surveys offer. The paper is honest about its limitations, including potential incompleteness of the search. It also ships useful appendices with methodology details and acronym definitions. However, the survey's value rests on the credibility of its systematic selection process and the accuracy of its quantitative synthesis; the internal inconsistencies in the screening counts directly undermine that credibility.

major comments (3)
  1. [Section 5.1 and Table 1] The screening counts are internally inconsistent, which prevents an independent reader from reconstructing the final pool. The text states that 59 papers were found via the database search and 3 additional papers were included, totaling 62 initial papers, while Table 1 reports a total of 76 initial papers (28+10+36+2). After abstract screening, the text says 31 articles met the inclusion criteria, but Table 1 sums to 30. Finally, the text says three 'other' papers were added, while Table 1 lists only two in the 'Others' row. Because every distribution, trend, and comparison in the paper is based on the final 26-paper pool, the selection procedure must be reproducible. Please correct the counts and provide a detailed flow diagram (e.g., PRISMA-style) that reconciles the number of papers at each stage.
  2. [Section 5.2.2 and Table 3] The claim that MoP and KEBLM 'overshadow the lower performing CPK and DAKI frameworks in all instances' is not fully supported by Table 3. Several cells are marked '/', indicating that results were not reported for those model/dataset combinations (e.g., MoP on NCBI, DAKI on HoC/PubMedQA/BioASQ7b, CPK on HoC/PubMedQA/BioASQ7b). Additionally, the caption states that baseline scores are taken from the original papers if given, and otherwise from MoP results, which mixes baseline sources across models and datasets; this makes cross-paper 'improvement' comparisons difficult to interpret. Please either restrict the comparison to settings where all methods have complete results on shared baselines or clearly qualify the comparison as partial and indicative rather than head-to-head.
  3. [Section 6 and Limitations] The survey's search was concluded in January 2024, yet the text cites 2024 publications such as Tian et al. (2024) and Vladika et al. (2024) that are not part of the final 26-paper pool. The Limitations section acknowledges that relevant work may have been overlooked, but the citation of these 2024 papers as evidence (e.g., for the continued use of MoP) without clarifying that they were not included in the systematic pool is confusing. Please state explicitly whether such post-search works are used only as contextual references or whether they should be part of the review, and if the latter, update the search cutoff accordingly.
minor comments (6)
  1. [Section 5.2.1] There is a typo in the Domain Analysis paragraph: 'common-secommon-sensding' should read 'common-sense reasoning' or similar.
  2. [Throughout] The terminology is inconsistent: the text uses both 'closed-domain' and 'close-domain'; please standardize to 'closed-domain'.
  3. [Table 2] Several rows in Table 2 have no nickname assigned (shown as '/'), which makes it awkward to refer to those papers. Consider assigning a short name to every included paper for readability.
  4. [Section 3.2] The sentence 'Another popular type of efficient adaptation we'd like to mention for completeness is low-rank adaptation or LoRA' is informal for a survey; consider rephrasing to a more neutral style.
  5. [Section 2.2] The phrase 'the most promising by the authors' is ambiguous; it would be clearer to state which authors (Colon-Hernandez et al. or Wei et al.) and which specific category they deemed most promising.
  6. [Appendix A.1] The exclusion criteria list does not mention the 'three additional papers' inclusion criterion; please document how these papers were selected and how they are distinguished from database results in the flow.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation found; the survey's taxonomy and performance comparison rest on external papers' reported results, with only minor non-load-bearing self-citations.

full rationale

The paper is a systematic literature review, so its central outputs are aggregations: the taxonomy of adapter types, the quantitative distributions, the qualitative themes, and the biomedical performance comparison in Table 3 are compiled from the reported results of the 26 included papers. No equation or construction in the paper defines a predicted quantity in terms of its own input, and no fitted parameter is renamed as a prediction. The claim that prior surveys miss adapter-based KELMs is a coverage assertion supported by the cited prior surveys (Colon-Hernandez et al., 2021; Wei et al., 2021), not a conclusion derived from the present paper's own assumptions. The only in-text self-citation of note is the sentence: "MoP, in particular, is being continually used for biomedical knowledge enhancement, even in 2024 (Vladika. et al., 2024)." This supports an ancillary remark in the recommendation discussion and is not load-bearing: the recommendation of MoP and KEBLM is grounded in the performance numbers in Table 3, which come from Meng et al. (2021) and Lai et al. (2023), not from the self-citation. The disclosure that the SLR is a reworked and updated version of a master's thesis by one of the authors is transparent provenance rather than an imported authority. The screening-count inconsistency between Section 5.1 (62 initial, 31 abstracts, 3 'other' papers) and Table 1 (76 initial, 30 abstracts, 2 'other' papers) is a reproducibility defect in the selection process, but it does not make any result equal to its input by construction; it is a rigor issue, not a circularity issue. No circular step meeting the quote-and-reduction standard is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey rests on the representativeness of the selected literature, the comparability of reported performance numbers, and the usefulness of the chosen taxonomic categories. No free parameters or invented entities are introduced.

assumptions (3)
  • domain assumption The search string and database selection identify the relevant literature on adapter-based KELMs.
    Section 4 inclusion criteria; authors acknowledge possible omissions in Limitations.
  • domain assumption Reported performance numbers from different original papers are comparable when combined into one table.
    Table 3 combines scores from multiple papers with different baselines and significance reporting; the comparison is not standardized.
  • domain assumption The taxonomic categories (adapter types, domains, tasks) are the appropriate organizing dimensions for the field.
    Categories are inherited from prior surveys and AdapterHub (Sections 2.2 and 3.2); the survey does not independently validate them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey." pith.science (2026). https://pith.science/paper/LGYE6W65

@misc{pith2026241116403,
  author       = {Pith},
  title        = {Pith review of: Adapter-based Approaches to Knowledge-enhanced Language Models -- A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LGYE6W65}},
  note         = {Machine review of arXiv:2411.16403}
}
read the original abstract

Knowledge-enhanced language models (KELMs) have emerged as promising tools to bridge the gap between large-scale language models and domain-specific knowledge. KELMs can achieve higher factual accuracy and mitigate hallucinations by leveraging knowledge graphs (KGs). They are frequently combined with adapter modules to reduce the computational load and risk of catastrophic forgetting. In this paper, we conduct a systematic literature review (SLR) on adapter-based approaches to KELMs. We provide a structured overview of existing methodologies in the field through quantitative and qualitative analysis and explore the strengths and potential shortcomings of individual approaches. We show that general knowledge and domain-specific approaches have been frequently explored along with various adapter architectures and downstream tasks. We particularly focused on the popular biomedical domain, where we provided an insightful performance comparison of existing KELMs. We outline the main trends and propose promising future directions.

Figures

Figures reproduced from arXiv: 2411.16403 by the authors.

Figure 1
Figure 1. Illustration of a standard fine-tuning versus a knowledge enhancement process. In the example, knowledge from a [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Location of the adapter module in a transformer [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Yearly distribution of publications Adapter Type Distribution Next, we evaluate the popularity and variety of adapter types used across the papers ( [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Distribution of adapter types used in the papers [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 69 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...

  2. [2]

    Alabi, J., Mosbach, M., Eyal, M., Klakow, D., and Geva, M. (2024). The hidden space of transformer language adapters. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , pages 6588--6607. Association for Computational Linguistics

  3. [3]

    P., Stork, L., and Groth, P

    Allen, B. P., Stork, L., and Groth, P. (2023). Knowledge Engineering Using Large Language Models . Transactions on Graph Data and Knowledge , 1(1):3:1--3:19

  4. [4]

    Auer, S., Bizer, C., Kobilarov, G., Lehmann, J., Cyganiak, R., and Ives, Z. (2007). Dbpedia: A nucleus for a web of open data. In The Semantic Web , pages 722--735, Berlin, Heidelberg. Springer Berlin Heidelberg

  5. [5]

    R., and Hinton, G

    Ba, J., Kiros, J. R., and Hinton, G. E. (2016). Layer normalization. ArXiv , abs/1607.06450

  6. [6]

    Baker, S., Silins, I., Guo, Y., Ali, I., Högberg, J., Stenius, U., and Korhonen, A. (2015). Automatic semantic classification of scientific literature according to the hallmarks of cancer . Bioinformatics , 32(3):432--440

  7. [7]

    and Firat, O

    Bapna, A. and Firat, O. (2019). Simple, scalable adaptation for neural machine translation. In Proceedings of the 2019 Conference on EMNLP-IJCNLP , pages 1538--1548, Hong Kong, China. Association for Computational Linguistics

  8. [8]

    Beltagy, I., Lo, K., and Cohan, A. (2019). Scibert: A pretrained language model for scientific text. In Conference on Empirical Methods in Natural Language Processing

Show all 81 references
  1. [9]

    Bodenreider, O. (2004). The unified medical language system (umls): integrating biomedical terminology. Nucleic acids research , 32

  2. [10]

    Budzianowski, P., Wen, T.-H., Tseng, B.-H., Casanueva, I., Ultes, S., Ramadan, O., and Ga s i \'c , M. (2018). M ulti WOZ - a large-scale multi-domain W izard-of- O z dataset for task-oriented dialogue modelling. In Proceedings of the 2018 Conference on Empirical Methods in Na...

  3. [11]

    and Knight, K

    Cai, S. and Knight, K. (2013). S match: an evaluation metric for semantic feature structures. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages 748--752, Sofia, Bulgaria. ACL

  4. [12]

    Chronopoulou, A., Peters, M., and Dodge, J. (2022). Efficient hierarchical domain adaptation for pretrained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , page...

  5. [13]

    Chronopoulou, A., Peters, M., Fraser, A., and Dodge, J. (2023). A dapter S oup: Weight averaging to improve generalization of pretrained language models. In Findings of the Association for Computational Linguistics: EACL 2023 , pages 2054--2063, Dubrovnik, Croatia. Association...

  6. [14]

    Cimini, G., Gabrielli, A., and Labini, F. (2014). The scientific competitiveness of nations. PloS one , 9

  7. [15]

    B., Huggins, M., and Breazeal, C

    Colon-Hernandez, P., Havasi, C., Alonso, J. B., Huggins, M., and Breazeal, C. (2021). Combining pre-trained language models and structured knowledge. ArXiv , abs/2101.12294

  8. [16]

    Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314

  9. [17]

    Diao, S., Xu, T., Xu, R., Wang, J., and Zhang, T. (2023). Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models ' memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages 5113--5...

  10. [18]

    Dinan, E., Roller, S., Shuster, K., Fan, A., Auli, M., and Weston, J. (2018). Wizard of wikipedia: Knowledge-powered conversational agents. ArXiv , abs/1811.01241

  11. [19]

    I., Leaman, R., and Lu, Z

    Dogan, R. I., Leaman, R., and Lu, Z. (2014). Ncbi disease corpus: A resource for disease name recognition and concept normalization. Journal of biomedical informatics , 47:1--10

  12. [20]

    Elsahar, H. (2017). T-Rex : A Large Scale Alignment of Natural Language with Knowledge Base Triples [NIF SAMPLE]

  13. [21]

    Emelin, D., Bonadiman, D., Alqahtani, S., Zhang, Y., and Mansour, S. (2022). Injecting domain knowledge in language models for task-oriented dialogue systems. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages 11962--11974. Associ...

  14. [22]

    Fichtl, A. (2024). Evaluating adapter-based knowledge-enhanced language models in the biomedical domain. Master's thesis, Technical University of Munich, Munich, Germany

  15. [23]

    R., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H

    Gu, Y., Tinn, R., Cheng, H., Lucas, M. R., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H. (2020). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH) , 3:1 -- 23

  16. [24]

    and Guo, Y

    Guo, Q. and Guo, Y. (2022). Lexicon enhanced chinese named entity recognition with pointer network. Neural Computing and Applications

  17. [25]

    Han, W., Pang, B., and Wu, Y. N. (2021). Robust transfer learning with pretrained language models through adapters. ArXiv , abs/2108.02340

  18. [26]

    He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G. (2021a). Towards a unified view of parameter-efficient transfer learning. ArXiv , abs/2110.04366

  19. [27]

    He, R., Liu, L., Ye, H., Tan, Q., Ding, B., Cheng, L., Low, J.-W., Bing, L., and Si, L. (2021b). On the effectiveness of adapter-based tuning for pretrained language model adaptation

  20. [28]

    He, Y., Zhu, Z., Zhang, Y., Chen, Q., and Caverlee, J. (2020). I nfusing D isease K nowledge into BERT for H ealth Q uestion A nswering, M edical I nference and D isease N ame R ecognition. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processi...

  21. [29]

    Hogan, A., Blomqvist, E., Cochez, M., d’Amato, C., de Melo, G., Guti \'e rrez, C., Kirrane, S., Gayo, J. E. L., Navigli, R., Neumaier, S., Ngomo, A.-C. N., Polleres, A., Rashid, S. M., Rula, A., Schmelzeisen, L., Sequeda, J., Staab, S., and Zimmermann, A. (2020). Knowledge gra...

  22. [30]

    Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019). Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning

  23. [31]

    J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

    Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). Lo RA : Low-rank adaptation of large language models. In International Conference on Learning Representations

  24. [32]

    Hu, L., Liu, Z., Zhao, Z., Hou, L., Nie, L., and Li, J. (2023). A survey of knowledge enhanced pre-trained language models

  25. [33]

    Huang , L., Yu , W., Ma , W., Zhong , W., Feng , Z., Wang , H., Chen , Q., Peng , W., Feng , X., Qin , B., and Liu , T. (2023). A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions . arXiv e-prints , page arXiv:2311.05232

  26. [34]

    Hung, C.-C., Lange, L., and Str \"o tgen, J. (2023). TADA : Efficient task-agnostic domain adaptation for transformers. In Findings of the Association for Computational Linguistics: ACL 2023 , pages 487--503, Toronto, Canada. Association for Computational Linguistics

  27. [35]

    Hung, C.-C., Lauscher, A., Ponzetto, S., and Glava s , G. (2022). DS - TOD : Efficient domain specialization for task-oriented dialog. In Findings of the Association for Computational Linguistics: ACL 2022 , pages 891--904. Association for Computational Linguistics

  28. [36]

    Ji, S., Pan, S., Cambria, E., Marttinen, P., and Yu, P. S. (2020). A survey on knowledge graphs: Representation, acquisition, and applications. IEEE Transactions on Neural Networks and Learning Systems , 33:494--514

  29. [37]

    Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X. (2019). P ub M ed QA : A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on EMNLP-IJCNLP , pages 2567--2577, Hong Kong, China. Association for Computational Linguistics

  30. [38]

    K r J rgensen, R., Hartmann, M., Dai, X., and Elliott, D. (2021). m DAPT : Multilingual domain adaptive pretraining in a single model. In Findings of the Association for Computational Linguistics: EMNLP 2021 , pages 3404--3418. Association for Computational Linguistics

  31. [39]

    C., Aliane, H., and Guessoum, A

    Khadir, A. C., Aliane, H., and Guessoum, A. (2021). Ontology learning: Grand tour and challenges. Computer Science Review , 39:100339

  32. [40]

    Kitchenham, B., Pearl Brereton , O., Budgen, D., Turner, M., Bailey, J., and Linkman, S. (2009). Systematic literature reviews in software engineering – a systematic literature review. Information and Software Technology , 51(1):7--15. Special Section - Most Cited Articles in ...

  33. [41]

    M., Zhai, C., and Ji, H

    Lai, T. M., Zhai, C., and Ji, H. (2023). Keblm: Knowledge-enhanced biomedical language models. Journal of Biomedical Informatics , 143:104392

  34. [42]

    Lauscher, A., Majewska, O., Ribeiro, L. F. R., Gurevych, I., Rozanov, N., and Glava s , G. (2020). Common sense or world knowledge? investigating adapter-based knowledge injection into pretrained transformers. In Proceedings of Deep Learning Inside Out (DeeLIO) , pages 43--49....

  35. [43]

    H., and Kang, J

    Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., and Kang, J. (2019). BioBERT: a pre-trained biomedical language representation model for biomedical text mining . Bioinformatics , 36(4):1234--1240

  36. [44]

    N., Chai Sim, K., Zhang, Y., Han, W., Strohman, T., and Beaufays, F

    Li, B., Hwang, D., Huo, Z., Bai, J., Prakash, G., Sainath, T. N., Chai Sim, K., Zhang, Y., Han, W., Strohman, T., and Beaufays, F. (2023). Efficient domain adaptation for speech foundation models. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Sig...

  37. [45]

    Liu, C., Zhang, S., Li, C., and Zhao, H. (2023). Cpk-adapter: Infusing medical knowledge into k-adapter with continuous prompt. In 2023 8th International Conference on Intelligent Computing and Signal Processing (ICSP) , pages 1017--1023, Los Alamitos, CA, USA. IEEE Computer Society

  38. [46]

    Lu, G., Yu, H., Yan, Z., and Xue, Y. (2023). Commonsense knowledge graph-based adapter for aspect-level sentiment classification. Neurocomputing , 534:67--76

  39. [47]

    Lu, Q., Dou, D., and Nguyen, T. H. (2021). Parameter-efficient domain knowledge integration from multiple sources for biomedical pre-trained language models. In Findings of the Association for Computational Linguistics: EMNLP 2021 , pages 3855--3865. Association for Computatio...

  40. [48]

    M., and Korhonen, A

    Majewska, O., Vuli \'c , I., Glava s , G., Ponti, E. M., and Korhonen, A. (2021). Verb knowledge injection for multilingual event processing. In Proceedings of the 59th Annual Meeting of the ACL-IJCNLP , pages 6952--6969. ACL

  41. [49]

    H., Shareghi, E., and Collier, N

    Meng, Z., Liu, F., Clark, T. H., Shareghi, E., and Collier, N. (2021). Mixture-of-partitions: Infusing large biomedical knowledge graphs into bert. ArXiv , abs/2109.04810

  42. [50]

    Moon, H., Park, C., Eo, S., Seo, J., and Lim, H. (2021). An empirical study on automatic post editing for neural machine translation. IEEE Access , 9:123754--123763

  43. [51]

    Nentidis, A., Bougiatiotis, K., Krithara, A., and Paliouras, G. (2020). Results of the seventh edition of the bioasq challenge. In Machine Learning and Knowledge Discovery in Databases , pages 553--568, Cham. Springer International Publishing

  44. [52]

    Nguyen-The, M., Lamghari, S., Bilodeau, G.-A., and Rockemann, J. (2023). Leveraging sentiment analysis knowledge to solve emotion detection tasks. In Pattern Recognition, Computer Vision, and Image Processing. ICPR 2022 International Workshops and Challenges , pages 405--416. ...

  45. [53]

    H., and Riedel, S

    Petroni, F., Rockt \"a schel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S. (2019). Language models as knowledge bases? ArXiv , abs/1909.01066

  46. [54]

    Pfeiffer, J., Kamath, A., R \"u ckl \'e , A., Cho, K., and Gurevych, I. (2020a). Adapterfusion: Non-destructive task composition for transfer learning. ArXiv , abs/2005.00247

  47. [55]

    Pfeiffer, J., R \"u ckl \'e , A., Poth, C., Kamath, A., Vuli \'c , I., Ruder, S., Cho, K., and Gurevych, I. (2020b). Adapterhub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrati...

  48. [56]

    Pfeiffer, J., Ruder, S., Vulić, I., and Ponti, E. M. (2024). Modular deep learning. Transcations of Machine Learning Research

  49. [57]

    Qian, Y., Gong, X., and Huang, H. (2022). Layer-wise fast adaptation for end-to-end multi-accent speech recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing , 30:2842--2853

  50. [58]

    Rebuffi, S.-A., Bilen, H., and Vedaldi, A. (2017). Learning multiple visual domains with residual adapters. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017 , pages 506--516

  51. [59]

    and Shivade, C

    Romanov, A. and Shivade, C. (2018). Lessons from natural language inference in the clinical domain. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 1586--1596, Brussels, Belgium. ACL

  52. [60]

    Rosset , C., Xiong , C., Phan , M., Song , X., Bennett , P., and Tiwary , S. (2020). Knowledge-Aware Language Model Pretraining . arXiv e-prints , page arXiv:2007.00655

  53. [61]

    Schabus, D., Skowron, M., and Trapp, M. (2017). One million posts: A data set of german online discussions. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval , page 1241–1244, New York, NY, USA. Association for C...

  54. [62]

    Schneider, P., Schopf, T., Vladika, J., Galkin, M., Simperl, E., and Matthes, F. (2022). A decade of knowledge graphs in natural language processing: A survey. In Proceedings of the 2nd AACL-IJCNLP , pages 601--614. Association for Computational Linguistics

  55. [63]

    Schuler, K. K. (2006). VerbNet: A Broad-Coverage, Comprehensive Verb Lexicon . PhD thesis, University of Pennsylvania

  56. [64]

    Shi, X., Yu, F., Lu, Y., Liang, Y., Feng, Q., Wang, D., Qian, Y., and Xie, L. (2021). The accented english speech recognition challenge 2020: Open datasets, tracks, baselines, results and methods. ICASSP 2021 - IEEE International Conference on Acoustics, Speech and Signal Proc...

  57. [65]

    Speer, R., Chin, J., and Havasi, C. (2017). Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence . AAAI Press

  58. [66]

    Stickland, A. C. and Murray, I. (2019). BERT and PAL s: Projected attention layers for efficient adaptation in multi-task learning. In Proceedings of the 36th International Conference on Machine Learning , pages 5986--5995. PMLR

  59. [67]

    Tian, S., Luo, Y., Xu, T., Yuan, C., Jiang, H., Wei, C., and Wang, X. (2024). KG -adapter: Enabling knowledge graph integration in large language models through parameter-efficient fine-tuning. In Findings of the Association for Computational Linguistics ACL 2024 , pages 3813-...

  60. [68]

    Tiwari, A., Manthena, M., Saha, S., Bhattacharyya, P., Dhar, M., and Tiwari, S. (2022). Dr. can see: Towards a multi-modal disease diagnosis virtual assistant. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management , page 1935–1944, New Y...

  61. [69]

    Tiwari, A., Saha, A., Saha, S., Bhattacharyya, P., and Dhar, M. (2023). Experience and evidence are the eyes of an excellent summarizer! towards knowledge infused multi-modal clinical conversation summarization. In Proceedings of the 32nd ACM International Conference on Inform...

  62. [70]

    M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A

    Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In NIPS

  63. [71]

    L., Mart \' nez Lorenzo, A

    Vasylenko, P., Huguet Cabot, P. L., Mart \' nez Lorenzo, A. C., and Navigli, R. (2023). Incorporating graph information in transformer-based AMR parsing. In Findings of the Association for Computational Linguistics: ACL 2023 , pages 1995--2011, Toronto, Canada. Association for...

  64. [72]

    Vladika., J., Fichtl., A., and Matthes., F. (2024). Diversifying knowledge enhancement of biomedical language models using adapter modules and knowledge graphs. In Proceedings of the 16th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART , pages...

  65. [73]

    Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2019). Glue: A multi-task benchmark and analysis platform for natural language understanding

  66. [74]

    Wang, R., Tang, D., Duan, N., Wei, Z., Huang, X., Ji, J., Cao, G., Jiang, D., and Zhou, M. (2020). K-adapter: Infusing knowledge into pre-trained models with adapters

  67. [75]

    Wei, X., Wang, S., Zhang, D., Bhatia, P., and Arnold, A. O. (2021). Knowledge enhanced pretrained language models: A compreshensive survey. ArXiv , abs/2110.08455

  68. [76]

    Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzm \'a n, F., Joulin, A., and Grave, E. (2020). CCN et: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference , pages 4003--4012

  69. [77]

    Wold, S. (2022). The effectiveness of masked language modeling and adapters for factual knowledge injection. In Proceedings of TextGraphs-16 , pages 54--59, Gyeongju, Republic of Korea. Association for Computational Linguistics

  70. [78]

    A., Tiwari, P., and Ananiadou, S

    Xie, Q., Bishop, J. A., Tiwari, P., and Ananiadou, S. (2022). Pre-trained language models with domain knowledge for biomedical extractive summarization. Knowledge-Based Systems , 252:109460

  71. [79]

    I., Madotto, A., Su, D., and Fung, P

    Xu, Y., Ishii, E., Cahyawijaya, S., Liu, Z., Winata, G. I., Madotto, A., Su, D., and Fung, P. (2022). Retrieval-free knowledge-grounded dialogue response generation with adapters. In Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Qu...

  72. [80]

    and Yang, Y

    Yu, S. and Yang, Y. (2023). A new feature fusion method based on pre-training model for sequence labeling. In 2023 6th International Conference on Data Storage and Data Engineering (DSDE) , pages 26--31

  73. [81]

    Zou, D., Zhang, X., Song, X., Yu, Y., Yang, Y., and Xi, K. (2022). Multiway bidirectional attention and external knowledge for multiple-choice reading comprehension. In 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages 694--699

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.