REVIEW 8 cited by
Towards Lifelong Learning of Large Language Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As the applications of large language models (LLMs) expand across diverse fields, the ability of these models to adapt to ongoing changes in data, tasks, and user preferences becomes crucial. Traditional training methods, relying on static datasets, are increasingly inadequate for coping with the dynamic nature of real-world information. Lifelong learning, also known as continual or incremental learning, addresses this challenge by enabling LLMs to learn continuously and adaptively over their operational lifetime, integrating new knowledge while retaining previously learned information and preventing catastrophic forgetting. This survey delves into the sophisticated landscape of lifelong learning, categorizing strategies into two primary groups: Internal Knowledge and External Knowledge. Internal Knowledge includes continual pretraining and continual finetuning, each enhancing the adaptability of LLMs in various scenarios. External Knowledge encompasses retrieval-based and tool-based lifelong learning, leveraging external data sources and computational tools to extend the model's capabilities without modifying core parameters. The key contributions of our survey are: (1) Introducing a novel taxonomy categorizing the extensive literature of lifelong learning into 12 scenarios; (2) Identifying common techniques across all lifelong learning scenarios and classifying existing literature into various technique groups within each scenario; (3) Highlighting emerging techniques such as model expansion and data selection, which were less explored in the pre-LLM era. Through a detailed examination of these groups and their respective categories, this survey aims to enhance the adaptability, reliability, and overall performance of LLMs in real-world applications.
Forward citations
Cited by 8 Pith papers
-
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
The paper argues that LLMs should be evaluated and built for meta-cultural competence rather than static knowledge of specific cultures, and gives a first, illustrative measurement of one component.
-
Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning
Autoregressive DNA language models fine-tuned jointly on classification, text generation, and image generation achieve strong benchmark results and open-ended cross-modal genomic tasks.
-
Learning Model Successors
A meta-learner that transforms a model for difficulty level k into a model for level k+1 can generalize far beyond its training range, demonstrated on a bracket-matching task.
-
Continual Gradient Low-Rank Projection Fine-Tuning for LLMs
GORP jointly trains LoRA and full-rank parameters inside a low-rank gradient subspace built from Adam first moments, reporting higher average accuracy and lower forgetting than O-LoRA and N-LoRA on LLM continual learn...
-
Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
BranchLoRA reduces catastrophic forgetting in multimodal continual instruction tuning by using a shared LoRA matrix, task-specific branches, frozen experts, and learned task keys.
-
Facilitating Video Story Interaction with Multi-Agent Collaborative System
A multi-agent system with VLM and RAG lets users talk with stage-aware Harry Potter characters and customize scenes, with a user study reporting enhanced engagement.
-
What Causes Knowledge Loss in Multilingual Language Models?
In sequential multilingual fine-tuning, non-Latin-script languages are forgotten more severely than Latin-script ones, and separating LoRA adapters per language reduces forgetting.
-
Infinite Video Understanding
The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.
Discussion (0). Continue with ORCID to comment.