REVIEW 3 cited by
Understanding Catastrophic Forgetting and Remembering in Continual Learning with Optimal Relevance Mapping
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Catastrophic forgetting in neural networks is a significant problem for continual learning. A majority of the current methods replay previous data during training, which violates the constraints of an ideal continual learning system. Additionally, current approaches that deal with forgetting ignore the problem of catastrophic remembering, i.e. the worsening ability to discriminate between data from different tasks. In our work, we introduce Relevance Mapping Networks (RMNs) which are inspired by the Optimal Overlap Hypothesis. The mappings reflects the relevance of the weights for the task at hand by assigning large weights to essential parameters. We show that RMNs learn an optimized representational overlap that overcomes the twin problem of catastrophic forgetting and remembering. Our approach achieves state-of-the-art performance across all common continual learning datasets, even significantly outperforming data replay methods while not violating the constraints for an ideal continual learning system. Moreover, RMNs retain the ability to detect data from new tasks in an unsupervised manner, thus proving their resilience against catastrophic remembering.
Forward citations
Cited by 3 Pith papers
-
Learning to Remember, Learn, and Forget in Attention-Based Models
Palimpsa adds a per-slot importance/precision state to gated linear attention, letting a fixed-size memory forget stale information and protect important information, and recovers Mamba2 as a high-forgetting limit.
-
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
A data composition method that aligns SFT data proportions with a model's detected domain knowledge distribution and dynamically reweights domains by learnable potential improves multi-domain performance versus unifor...
-
Modality-Incremental Learning with Disjoint Relevance Mapping Networks for Image-based Semantic Segmentation
Disjoint Relevance Mapping Networks, which forbid weight sharing across sensor modalities, reduce forgetting in incremental semantic segmentation but only slightly outperform the shared-weight RMN baseline.
Discussion (0). Continue with ORCID to comment.