REVIEW 10 cited by
Overcoming catastrophic forgetting in neural networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to learn tasks in a sequential fashion is crucial to the development of artificial intelligence. Neural networks are not, in general, capable of this and it has been widely thought that catastrophic forgetting is an inevitable feature of connectionist models. We show that it is possible to overcome this limitation and train networks that can maintain expertise on tasks which they have not experienced for a long time. Our approach remembers old tasks by selectively slowing down learning on the weights important for those tasks. We demonstrate our approach is scalable and effective by solving a set of classification tasks based on the MNIST hand written digit dataset and by learning several Atari 2600 games sequentially.
Forward citations
Cited by 10 Pith papers
-
TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners
OOD detection in continual learning degrades through task-dependent logit-scale drift and feature-space crowding; a post-hoc per-task energy calibration recovers most of that loss for energy-based detectors.
-
LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
Adapting a text-to-image generator with task-specific LoRA adapters and filtering samples by the model's own confidence improves synthetic replay in continual vision-language learning.
-
Temporal Information Retrieval via Time-Specifier Model Merging
TSM trains one retriever per time specifier and merges them by parameter averaging, improving temporal retrieval while maintaining non-temporal retrieval.
-
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.
-
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness
IAM interpolates between an original model and a shadow model to score each sample's unlearning completeness, achieving top AUC for exact unlearning and top correlation for approximate unlearning, and exposing under- ...
-
Memoir: Should a Model Write to Its Memory While It Thinks?
Writing to fast memory during pondering slows associative-recall learning at a fixed budget, but does not reduce final performance once training is long enough.
-
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
In a two-phase Terminal-Bench evaluation, only regression-aware RELAI-VCL compounded optimization gains, reaching the highest pass rate at every stage.
-
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
Quantization tends to degrade LLM fairness and safety—more in non-English tasks—and preserving top sensitivity-ranked weights in FP16 mostly mitigates the loss.
-
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
Quantized (INT8/INT4) LLMs can outperform FP16 in later-task forward accuracy and retention during continual learning, though single-seed runs leave the effect unquantified.
-
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
A continual multimodal misinformation detector that uses Dirichlet process-based expert expansion to curb forgetting and a neural-ODE dynamics model to anticipate evolving fake-news distributions.
Discussion (0). Sign in to comment.