Pith. sign in

REVIEW 10 cited by

Overcoming catastrophic forgetting in neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.00796 v2 pith:36TDEJBL submitted 2016-12-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords tasksnetworksapproachcatastrophicforgettinglearningneuralability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to learn tasks in a sequential fashion is crucial to the development of artificial intelligence. Neural networks are not, in general, capable of this and it has been widely thought that catastrophic forgetting is an inevitable feature of connectionist models. We show that it is possible to overcome this limitation and train networks that can maintain expertise on tasks which they have not experienced for a long time. Our approach remembers old tasks by selectively slowing down learning on the weights important for those tasks. We demonstrate our approach is scalable and effective by solving a set of classification tasks based on the MNIST hand written digit dataset and by learning several Atari 2600 games sequentially.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OOD detection in continual learning degrades through task-dependent logit-scale drift and feature-space crowding; a post-hoc per-task energy calibration recovers most of that loss for energy-based detectors.

  2. LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Adapting a text-to-image generator with task-specific LoRA adapters and filtering samples by the model's own confidence improves synthetic replay in continual vision-language learning.

  3. Temporal Information Retrieval via Time-Specifier Model Merging

    cs.IR 2025-07 conditional novelty 6.0 of 10

    TSM trains one retriever per time specifier and merges them by parameter averaging, improving temporal retrieval while maintaining non-temporal retrieval.

  4. Universal Music Representations? Evaluating Foundation Models on World Music Corpora

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.

  5. Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness

    cs.LG 2025-06 conditional novelty 6.0 of 10

    IAM interpolates between an original model and a shadow model to score each sample's unlearning completeness, achieving top AUC for exact unlearning and top correlation for approximate unlearning, and exposing under- ...

  6. Memoir: Should a Model Write to Its Memory While It Thinks?

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Writing to fast memory during pondering slows associative-recall learning at a fixed budget, but does not reduce final performance once training is long enough.

  7. Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

    cs.AI 2026-07 conditional novelty 5.0 of 10

    In a two-phase Terminal-Bench evaluation, only regression-aware RELAI-VCL compounded optimization gains, reaching the highest pass rate at every stage.

  8. Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Quantization tends to degrade LLM fairness and safety—more in non-English tasks—and preserving top sensitivity-ranked weights in FP16 mostly mitigates the loss.

  9. When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models

    cs.LG 2025-12 conditional novelty 5.0 of 10

    Quantized (INT8/INT4) LLMs can outperform FP16 in later-task forward accuracy and retention during continual learning, though single-seed runs leave the effect unquantified.

  10. Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A continual multimodal misinformation detector that uses Dirichlet process-based expert expansion to curb forgetting and a neural-ODE dynamics model to anticipate evolving fake-news distributions.

Pith tools