{"id":"dfc88efe-6eaa-4bc7-969d-58d72ca8e3b9","arxiv_id":"2411.16403","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of adapter-based knowledge-enhanced language models, covering 26 papers, popular adapter types, and biomedical performance comparisons.","lead":"This paper is a systematic literature review of methods that plug knowledge graphs into language models through lightweight adapter modules. It maps the field, compares biomedical results, and identifies trends such as the dominance of Pfeiffer and Houlsby adapters.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Screening counts in Table 1 contradict Section 5.1 (76 vs 62 initial; 30 vs 31 abstracts; 2 vs 3 'other' papers), making the SLR's selection process non-reproducible; this is the load-bearing weakness to fix.","rationale":"The reader's weakest_assumption focused on representativeness and completeness of the literature search, which is a genuine limitation the authors themselves acknowledge. My stress-test identifies a different, more concrete weakness: the reported screening counts are internally inconsistent. This matters because a systematic review's central claim depends on a reproducible selection process; if the numbers do not add up, the trustworthiness of the 26-paper pool is compromised. The reader did mention inconsistent screening counts in the rationale, so my concern partially overlaps, but I do not treat external search coverage as the primary issue. The paper still has value as a qualitative map and the broad trends are plausible, so the condition attached to acceptance is appropriate: correct the screening numbers and make the selection fully auditable. No change to the reader's CONDITIONAL verdict is needed.","tokens_in":15505,"tokens_out":6145,"duration_ms":59531,"concrete_test":"Reconstruct the PRISMA-style flow exactly as specified in Section 4 and Appendix A.1: re-run the stated query on IEEE Xplore, ACM Digital Library, and ACL Anthology as of January 2024, and record initial hits, abstracts accepted, and full texts accepted. Compare each count to Table 1 and to the prose in Section 5.1. Then audit Table 2 by assigning each of the 26 papers to its actual publication venue and screening stage, verifying whether the 'Others' column should be 2 or 3 and whether papers such as K-Adapter and MoP are classified correctly. If the correct counts match Table 1, the discrepancy is a typographical error and the concern is cosmetic; if they match the prose, Table 1 and all downstream totals must be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a systematic literature review: the value of the survey rests on a transparent, reproducible selection procedure. That procedure does not currently reconcile. Section 5.1 states that 59 papers were found via the database search and 3 additional papers were included, for 62 initial papers; Table 1 reports 76 initial papers (28+10+36+2). After abstract screening, the text reports 31 articles meeting inclusion criteria, while Table 1 sums to 30 abstracts. The text says three 'other' papers were added, but Table 1 lists only two in the 'Others' row. Because the final pool of 26 papers is the basis for every distribution, trend, and performance comparison in the paper, an internally inconsistent screening count means the reader cannot verify how those 26 papers were selected. The authors' own Limitations section already acknowledges potential incompleteness, so the more pressing issue is not coverage but internal consistency: the methodology section must allow an independent auditor to reconstruct the pool. Without corrected counts, the survey's claim to be systematic is weakened, even though the individual paper summaries and qualitative observations may still be useful.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review (SLR) of adapter-based approaches to knowledge-enhanced language models (KELMs). The authors identify a gap in prior surveys (Colon-Hernandez et al., 2021; Wei et al., 2021) regarding adapter-based methods and aim to fill it by reviewing 26 papers from ACM, ACL, IEEE, and other sources. The paper categorizes adapter architectures (Houlsby, Pfeiffer, Bapna and Firat, K-Adapter, and unique variants), domains (open vs. closed, with biomedical as the most frequent closed domain), and downstream tasks. It provides quantitative trends (yearly growth, adapter-type distribution, domain distribution) and a qualitative synthesis of general, linguistic, domain-specific, and biomedical knowledge injection approaches. A particular contribution is the biomedical performance comparison in Table 3, which reports accuracy/F1 on five tasks across three base models enhanced with MoP, KEBLM, DAKI, and CPK. The paper concludes with current trends and future directions. The final pool of 26 papers, their categorizations, and the performance comparison form the empirical basis of the survey.","tokens_in":15631,"tokens_out":2857,"duration_ms":27843,"significance":"If the review is reproducible and the reported comparisons are sound, the paper provides a valuable structured entry point to a nascent and fast-growing subfield. The taxonomy of adapter types and the domain/task categorization are useful organizational contributions, and the biomedical performance comparison, though focused, is more specific than what prior KELM surveys offer. The paper is honest about its limitations, including potential incompleteness of the search. It also ships useful appendices with methodology details and acronym definitions. However, the survey's value rests on the credibility of its systematic selection process and the accuracy of its quantitative synthesis; the internal inconsistencies in the screening counts directly undermine that credibility.","major_comments":[{"comment":"The screening counts are internally inconsistent, which prevents an independent reader from reconstructing the final pool. The text states that 59 papers were found via the database search and 3 additional papers were included, totaling 62 initial papers, while Table 1 reports a total of 76 initial papers (28+10+36+2). After abstract screening, the text says 31 articles met the inclusion criteria, but Table 1 sums to 30. Finally, the text says three 'other' papers were added, while Table 1 lists only two in the 'Others' row. Because every distribution, trend, and comparison in the paper is based on the final 26-paper pool, the selection procedure must be reproducible. Please correct the counts and provide a detailed flow diagram (e.g., PRISMA-style) that reconciles the number of papers at each stage.","section":"Section 5.1 and Table 1"},{"comment":"The claim that MoP and KEBLM 'overshadow the lower performing CPK and DAKI frameworks in all instances' is not fully supported by Table 3. Several cells are marked '/', indicating that results were not reported for those model/dataset combinations (e.g., MoP on NCBI, DAKI on HoC/PubMedQA/BioASQ7b, CPK on HoC/PubMedQA/BioASQ7b). Additionally, the caption states that baseline scores are taken from the original papers if given, and otherwise from MoP results, which mixes baseline sources across models and datasets; this makes cross-paper 'improvement' comparisons difficult to interpret. Please either restrict the comparison to settings where all methods have complete results on shared baselines or clearly qualify the comparison as partial and indicative rather than head-to-head.","section":"Section 5.2.2 and Table 3"},{"comment":"The survey's search was concluded in January 2024, yet the text cites 2024 publications such as Tian et al. (2024) and Vladika et al. (2024) that are not part of the final 26-paper pool. The Limitations section acknowledges that relevant work may have been overlooked, but the citation of these 2024 papers as evidence (e.g., for the continued use of MoP) without clarifying that they were not included in the systematic pool is confusing. Please state explicitly whether such post-search works are used only as contextual references or whether they should be part of the review, and if the latter, update the search cutoff accordingly.","section":"Section 6 and Limitations"}],"minor_comments":[{"comment":"There is a typo in the Domain Analysis paragraph: 'common-secommon-sensding' should read 'common-sense reasoning' or similar.","section":"Section 5.2.1"},{"comment":"The terminology is inconsistent: the text uses both 'closed-domain' and 'close-domain'; please standardize to 'closed-domain'.","section":"Throughout"},{"comment":"Several rows in Table 2 have no nickname assigned (shown as '/'), which makes it awkward to refer to those papers. Consider assigning a short name to every included paper for readability.","section":"Table 2"},{"comment":"The sentence 'Another popular type of efficient adaptation we'd like to mention for completeness is low-rank adaptation or LoRA' is informal for a survey; consider rephrasing to a more neutral style.","section":"Section 3.2"},{"comment":"The phrase 'the most promising by the authors' is ambiguous; it would be clearer to state which authors (Colon-Hernandez et al. or Wei et al.) and which specific category they deemed most promising.","section":"Section 2.2"},{"comment":"The exclusion criteria list does not mention the 'three additional papers' inclusion criterion; please document how these papers were selected and how they are distinguished from database results in the flow.","section":"Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reworked version of a master's thesis, which is not inherently problematic. However, the authors should be asked to address the screening count inconsistencies rigorously before the review can be considered systematic. The self-citation (Vladika et al., 2024) used to support the claim that MoP is 'continually used' should be de-emphasized or supplemented by independent citations. Overall, the paper's contribution is primarily organizational, so the methodology must be beyond reproach for the survey to be published as an SLR."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper fills a real gap: prior KELM surveys (Colon-Hernandez, Wei et al.) barely touch adapter-based methods, and this one gives you a structured table of 26 papers, a taxonomy by adapter type (Houlsby, Pfeiffer, K-Adapter, Bapna/Firat), domain, and tasks, plus a biomedical comparison. That's genuinely useful if you're entering this area or looking for a starting point.\n\nThe main problem is exactly what the stress-test flagged: the screening counts don't reconcile. Section 5.1 says 59 papers from databases plus 3 others = 62 initial; Table 1 says 76 (28 IEEE, 10 ACM, 36 ACL, 2 others). Abstract screening: text says 31, table sums to 30. And the 'three additional papers' in the text appear as two in the table. For an SLR, the selection process is the load-bearing part. If you can't reconstruct how the 26 papers were chosen, the 'systematic' claim weakens considerably. This isn't a cosmetic typo; it's the core method.\n\nThat said, the broad trends they report are plausible and consistent with the table: Pfeiffer and Houlsby dominate, biomedical is the main closed-domain area, interest is growing. The limitations section already concedes that coverage may be incomplete due to the restricted database list (ACM, ACL, IEEE only, no arXiv) and the January 2024 cutoff, so we know what we're getting. The performance comparison in Table 3 is a bit loose—baselines and significance flags come from different papers—so treat those numbers as suggestive, not a proper benchmark.\n\nOne more thing: the claim that MoP is 'continually used' rests on a citation to the authors' own work (Vladika et al., 2024). That's a minor self-citation, not a fatal flaw, but it's worth noting.\n\nBottom line: this is a well-meaning, useful survey with a real methodological wrinkle. It deserves peer review because the fixes are tractable: reconcile the counts, add a PRISMA-style flow diagram, and be explicit about how the initial pool maps to the final 26. After that, it would be a solid reference. I'd cite it for the map, not for the numbers.","headline":"Useful map of adapter-based KELMs, but the screening counts don't add up — fix that before trusting the 'systematic' claims.","tokens_in":16222,"tokens_out":2776,"would_cite":true,"duration_ms":24365,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper is a systematic literature review that assembles 26 adapter-based methods for injecting knowledge into language models, classifies them by architecture, domain, and task, and shows that adapter-based approaches consistently…","keywords":["knowledge-enhanced language models","adapter modules","knowledge graphs","systematic literature review","biomedical natural language processing","parameter-efficient fine-tuning","knowledge injection","survey"],"falsifier":"A reader could rerun the search with the same inclusion criteria plus additional databases such as arXiv and Scopus and broader query synonyms such as 'parameter-efficient' or 'prompt tuning'; if this expanded search surfaces a substantial body of adapter-based KELM papers absent from the 26, the review's claim to comprehensive coverage would be weakened.","tokens_in":109,"feed_emoji":"🧩","tokens_out":5691,"duration_ms":111284,"temperature":0.7,"pith_summary":"The paper is a systematic literature review of adapter-based knowledge-enhanced language models (KELMs): large language models that receive structured knowledge, usually from knowledge graphs, through small trainable adapter modules rather than full fine-tuning. It assembles 26 peer-reviewed papers from a structured search and classifies them by adapter architecture, domain scope, knowledge source, and downstream task. It finds that open-domain 'general knowledge' injection and closed-domain biomedical injection have both been actively explored, that the most common adapter architectures are the original bottleneck design and its simplified single-module variant, and that on shared biomedical benchmarks the MoP and KEBLM frameworks deliver the largest accuracy gains over base models. The contribution is a structured map of a young field, intended as an entry point for researchers.","feed_headline":"26 adapter methods for knowledge-enhanced language models, compared","feed_subtitle":"A systematic review finds biomedical models lead domain-specific work, with two frameworks delivering the largest accuracy gains.","key_machinery":"The central object is the adapter module itself: a small bottleneck feed-forward block inserted into a frozen transformer layer, adding only a few percent of new parameters and trained on external knowledge while the base model stays fixed. The review's machinery is a categorization scheme that sorts the 26 papers along four axes: adapter type (original bottleneck, single-adapter variant, K-Adapter plug-ins, and custom designs), domain scope (open vs. closed, with biomedical singled out), knowledge source (ConceptNet, DBpedia, UMLS, and others), and downstream task (question answering, named entity recognition, reading comprehension, dialogue, and more). This scheme is what turns a scattered set of papers into a comparable landscape, and it supports the paper's head-to-head biomedical performance table.","core_discovery":"On the paper's own terms, the discovery is that adapter-based knowledge enhancement is not a niche corner of the KELM literature but a recognizably distinct and rapidly growing approach with a clear methodological structure. The review shows that general-knowledge and domain-specific approaches have been frequently explored, identifies the single-adapter-bottleneck configuration as the predominant adapter type, and documents the biomedical domain as the most active closed-domain application area. Its quantitative comparison on five biomedical benchmarks (HoC, PubMedQA, BioASQ7b, MedNLI, and NCBI) reports that adapter-based KELMs consistently improve over their base models, with the MoP and KEBLM frameworks producing the strongest gains and the CPK and DAKI frameworks trailing. The paper's central claim is that prior surveys missed most of this adapter-based work, so this review fills that gap.","pith_inferences":["If the performance pattern holds beyond the five benchmarks, adapter-based knowledge injection may become the default low-cost route for domain adaptation in regulated fields, since adapters preserve the base model and reduce catastrophic forgetting.","The review's exclusion of LoRA/QLoRA from the core pool, despite noting them, hints that the boundary between 'adapter layers' and 'low-rank weight updates' is becoming blurred; a future survey might treat them under one parameter-efficient umbrella.","A testable extension is to run a controlled comparison of the four adapter types on a single set of knowledge-intensive tasks with identical base models and knowledge sources, since the paper's cross-paper comparison cannot fully control for training data and hyperparameters.","The biomedical dominance may reflect the availability of UMLS and similar curated knowledge graphs, so applying the same adapter recipe to legal or financial knowledge graphs could transfer the observed gains."],"forward_implications":["A new researcher can use the classification to locate the adapter architecture and benchmark most likely to suit a target domain.","Knowledge-intensive tasks such as question answering and named entity recognition are the established testbeds, while generative tasks like summarization and open-ended text generation remain largely unexplored.","In the biomedical domain, MoP and KEBLM are the recommended frameworks when the goal is maximising accuracy on the five shared benchmarks.","The roughly linear yearly increase in publications suggests the field will keep expanding with novel adapter architectures.","The balance between open-domain and closed-domain approaches, with biomedical dominant among closed domains, points to legal or financial document understanding as likely next applications."],"supporting_citations":[{"why":"Defines the bottleneck adapter module that the review treats as the field's starting point and sets the inclusion date of February 2019.","marker":"Houlsby et al. (2019)"},{"why":"Introduces AdapterFusion and the simplified single-adapter configuration that the review identifies as the most popular adapter type.","marker":"Pfeiffer et al. (2020a)"},{"why":"One of the two prior KELM surveys the paper claims largely missed adapter-based approaches, establishing the gap this review fills.","marker":"Colon-Hernandez et al. (2021)"},{"why":"The second prior KELM survey that the paper argues provides taxonomies but not systematic coverage of adapter-based methods.","marker":"Wei et al. (2021)"},{"why":"Introduces K-Adapter, one of the few adapter-based KELM works that prior surveys did cover and a distinct architecture class in the review.","marker":"Wang et al. (2020)"},{"why":"Presents the MoP framework, which the biomedical performance comparison identifies as one of the two strongest approaches.","marker":"Meng et al. (2021)"},{"why":"Presents the KEBLM framework, the other top performer in the review's biomedical benchmark comparison.","marker":"Lai et al. (2023)"},{"why":"Supplies the systematic literature review methodology that the paper follows for its search and selection process.","marker":"Kitchenham et al. (2009)"}],"fun_headline_variants":["Survey fills gap on adapter-based knowledge-enhanced LMs","Adapter-based KELMs outperform baselines on biomedical tasks","Systematic review maps adapter-based KELMs and their gains","The adapter route to knowledge-enhanced LMs, systematically reviewed"],"cache_read_input_tokens":18304,"weakest_assumption_plain":"The survey's map of the field is only as complete as its literature search, which restricted inclusion to English peer-reviewed papers in three bibliographic databases matching a fixed query and stopped in January 2024, so relevant work using other terms or venues could be missing.","fun_headline_variants_meta":{"raw":{"variants":["Survey fills gap on adapter-based knowledge-enhanced LMs","Adapter-based KELMs outperform baselines on biomedical tasks","Systematic review maps adapter-based KELMs and their gains","The adapter route to knowledge-enhanced LMs, systematically reviewed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001164,"raw_usage":{"total_tokens":4770,"prompt_tokens":852,"completion_tokens":3918,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":3851}},"tokens_in":468,"tokens_out":3918,"duration_ms":27717,"temperature":1.0,"reasoning_tokens":3851,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:09:17.022642+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could rerun the search with the same inclusion criteria plus additional databases such as arXiv and Scopus and broader query synonyms such as 'parameter-efficient' or 'prompt tuning'; if this expanded search surfaces a substantial body of adapter-based KELM papers absent from the 26, the review's claim to comprehensive coverage would be weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the bottleneck adapter module that the review treats as the field's starting point and sets the inclusion date of February 2019."},{"cited_title":"Combining pre-trained language models and structured knowledge","cited_arxiv_id":"2101.12294","evidence_quote":"One of the two prior KELM surveys the paper claims largely missed adapter-based approaches, establishing the gap this review fills."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces K-Adapter, one of the few adapter-based KELM works that prior surveys did cover and a distinct architecture class in the review."},{"cited_title":"Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT","cited_arxiv_id":"2109.04810","evidence_quote":"Presents the MoP framework, which the biomedical performance comparison identifies as one of the two strongest approaches."},{"cited_title":"M., Zhai, C., and Ji, H","cited_arxiv_id":null,"evidence_quote":"Presents the KEBLM framework, the other top performer in the review's biomedical benchmark comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the systematic literature review methodology that the paper follows for its search and selection process."}],"review_version":1}