Pith. sign in

REVIEW 31 cited by

Overcoming catastrophic forgetting in neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.00796 v2 pith:36TDEJBL submitted 2016-12-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords tasksnetworksapproachcatastrophicforgettinglearningneuralability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The ability to learn tasks in a sequential fashion is crucial to the development of artificial intelligence. Neural networks are not, in general, capable of this and it has been widely thought that catastrophic forgetting is an inevitable feature of connectionist models. We show that it is possible to overcome this limitation and train networks that can maintain expertise on tasks which they have not experienced for a long time. Our approach remembers old tasks by selectively slowing down learning on the weights important for those tasks. We demonstrate our approach is scalable and effective by solving a set of classification tasks based on the MNIST hand written digit dataset and by learning several Atari 2600 games sequentially.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Continual Learning with Maximally Interfered Retrieval

    cs.LG 2019-08 accept novelty 7.0 of 10

    Selecting replay samples by estimated loss increase after a virtual update improves online continual learning performance over random replay.

  2. An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A domain-adapted 8B LLM with agentic guideline retrieval matched UF Health oncologists' ratings on correctness, currency, and safety for colorectal cancer treatment plans in a 79-case blinded evaluation.

  3. In-Context Collapse in Vision-Language Models and How to Mitigate it?

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Many-shot in-context learning in vision-language models can collapse accuracy as demonstrations accumulate, and the failure is causally localized to the vision-language integration pathway, where a small adapter repairs it.

  4. TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OOD detection in continual learning degrades through task-dependent logit-scale drift and feature-space crowding; a post-hoc per-task energy calibration recovers most of that loss for energy-based detectors.

  5. LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Adapting a text-to-image generator with task-specific LoRA adapters and filtering samples by the model's own confidence improves synthetic replay in continual vision-language learning.

  6. Temporal Information Retrieval via Time-Specifier Model Merging

    cs.IR 2025-07 conditional novelty 6.0 of 10

    TSM trains one retriever per time specifier and merges them by parameter averaging, improving temporal retrieval while maintaining non-temporal retrieval.

  7. Universal Music Representations? Evaluating Foundation Models on World Music Corpora

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.

  8. Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness

    cs.LG 2025-06 conditional novelty 6.0 of 10

    IAM interpolates between an original model and a shadow model to score each sample's unlearning completeness, achieving top AUC for exact unlearning and top correlation for approximate unlearning, and exposing under- ...

  9. OOD Detection with immature Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Partially trained GLOW models match or outperform fully trained models for out-of-distribution image detection when scored by layer-wise gradient norms.

  10. Chained Tuning Leads to Biased Forgetting

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Fine-tuning a safety-tuned LLM on a capability task erases safety behavior more than the reverse order, and this forgetting is worse for specific groups such as Muslim people in the authors' tests.

  11. TinySubNets: An efficient and low capacity continual learning strategy

    cs.LG 2024-12 conditional novelty 6.0 of 10

    TinySubNets combines per-layer pruning, adaptive quantization, and KL-gated weight sharing to run continual image classification with much lower memory than PackNet, WSN, and Ada-QPacknet.

  12. MyTimeMachine: Personalized Facial Age Transformation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A personalized facial age transformation method that uses an adapter network on top of the SAM global aging model, trained with 10 to 50 photos of one person, to produce re-aged images that resemble that person's actu...

  13. Improved Cardinality Estimation by Learning Queries Containment Rates

    cs.DB 2019-08 conditional novelty 6.0 of 10

    Learned containment rates between query pairs, combined with a queries pool of known cardinalities, substantially improve cardinality estimates on multi-join queries.

  14. Interleaved Multitask Learning for Audio Source Separation with Independent Databases

    cs.SD 2019-08 conditional novelty 6.0 of 10

    An interleaved multitask training procedure for a shared-encoder source separation network enables training on independent per-source databases and yields SIR improvements over simultaneous multitask training.

  15. Developing LLM-based Multi-Agent Systems in Software Engineering: A Mixed-Method Experience Report

    cs.SE 2026-08 reject novelty 5.0 of 10

    A mixed-method experience report on LLM-based multi-agent frameworks for software engineering: broad feature coverage, weak monitoring support, and no clear quality winner on a README-summarization task, with incomple...

  16. Memoir: Should a Model Write to Its Memory While It Thinks?

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Writing to fast memory during pondering slows associative-recall learning at a fixed budget, but does not reduce final performance once training is long enough.

  17. Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

    cs.AI 2026-07 conditional novelty 5.0 of 10

    In a two-phase Terminal-Bench evaluation, only regression-aware RELAI-VCL compounded optimization gains, reaching the highest pass rate at every stage.

  18. Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Quantization tends to degrade LLM fairness and safety—more in non-English tasks—and preserving top sensitivity-ranked weights in FP16 mostly mitigates the loss.

  19. When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models

    cs.LG 2025-12 conditional novelty 5.0 of 10

    Quantized (INT8/INT4) LLMs can outperform FP16 in later-task forward accuracy and retention during continual learning, though single-seed runs leave the effect unquantified.

  20. Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A continual multimodal misinformation detector that uses Dirichlet process-based expert expansion to curb forgetting and a neural-ODE dynamics model to anticipate evolving fake-news distributions.

  21. CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

    cs.SD 2025-06 conditional novelty 5.0 of 10

    CultureMERT adapts the MERT-95M music model to Greek, Turkish and Indian traditions via two-stage continual pre-training, improving non-Western auto-tagging by 4.9% on average while keeping Western performance intact.

  22. What is the role of memorization in Continual Learning?

    cs.LG 2025-05 conditional novelty 5.0 of 10

    High-memorization training examples are forgotten fastest in class-incremental learning, and a cheap proxy based on learning iteration can guide buffer policies, favoring typical samples for small buffers and memorize...

  23. Chain-of-Model Learning for Language Model

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A nested Transformer with causally ordered hidden chains offers multiple sub-model sizes, chain-based expansion, and KV-cache sharing for faster prefilling.

  24. Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

    cs.LG 2025-02 reject novelty 5.0 of 10

    A reward-weighted flow matching loss, regularized by a Wasserstein-2 vector-field penalty, fine-tunes flow generators online with a controllable reward-diversity trade-off.

  25. Scaling Sequential Recommendation Models with Transformers

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Transformer-based sequential recommenders exhibit power-law and saturating NDCG scaling with model size and training interactions, enabling compute-aware model selection and effective pre-train/fine-tune transfer.

  26. MTSpark: Enabling Multi-Task Learning with Spiking Neural Networks for Generalist Agents

    cs.NE 2024-12 reject novelty 5.0 of 10

    MTSpark combines active dendrites and dueling in a deep spiking Q-network, reporting strong multi-task RL and classification scores, but the experimental setup may not fairly test the claimed continual-learning benefit.

  27. Targeting Negative Flips in Active Learning using Validation Sets

    cs.LG 2024-11 conditional novelty 5.0 of 10

    RoSE, a validation-set-based filter that restricts active learning acquisition functions to estimated negative flips, improves accuracy and/or reduces negative flip rates on several image benchmarks.

  28. The Future of Continual Learning in the Era of Foundation Models: Three Key Directions

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.

  29. Brain-Inspired Quantum Neural Architectures for Pattern Recognition: Integrating QSNN and QLSTM

    cs.ET 2025-05 reject novelty 4.0 of 10

    A hybrid QSNN and QLSTM model, trained in three phases, is claimed to beat classical and quantum baselines on imbalanced credit card fraud data, but the benchmark lacks matched controls and statistical rigor.

  30. Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning

    cs.RO 2024-12 conditional novelty 3.0 of 10

    A simulated mobile robot learns multi-step household instructions better when training is staged from short sub-goals to full instructions, but the supporting experiments lack quantitative comparison.

  31. Energy-Aware Deep Learning on Resource-Constrained Hardware

    cs.LG 2025-05 conditional novelty 1.0 of 10

    A survey of energy-aware deep learning methods for resource-constrained devices, covering energy-aware design, adaptive inference, on-device training, and scheduling on energy-harvesting systems.

Pith tools