Pith. sign in

REVIEW 2 cited by

KILM: Knowledge Injection into Encoder-Decoder Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.09170 v1 pith:ODRKRMDZ submitted 2023-02-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords knowledgemodelskilmlanguageparametersplmstasksencoder-decoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large pre-trained language models (PLMs) have been shown to retain implicit knowledge within their parameters. To enhance this implicit knowledge, we propose Knowledge Injection into Language Models (KILM), a novel approach that injects entity-related knowledge into encoder-decoder PLMs, via a generative knowledge infilling objective through continued pre-training. This is done without architectural modifications to the PLMs or adding additional parameters. Experimental results over a suite of knowledge-intensive tasks spanning numerous datasets show that KILM enables models to retain more knowledge and hallucinate less, while preserving their original performance on general NLU and NLG tasks. KILM also demonstrates improved zero-shot performances on tasks such as entity disambiguation, outperforming state-of-the-art models having 30x more parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor

    cs.CL 2025-01 conditional novelty 6.0 of 10

    K-COMP generates entity definitions and a compressed summary from retrieved medical passages, improving retrieval-augmented QA over baseline compressors on MedQuAD, MASH-QA, and BioASQ.

  2. CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A continual pre-training framework combining query-item joint training, in-context pre-training on related queries/items, and teacher-generated reading comprehension data improves LLM relevance modeling in commercial search.

Pith tools