Pith. sign in

REVIEW 2 major objections 6 minor 94 references

Flashbacks to Harmonize Stability and Plasticity in Continual Learning

T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Flashback Learning claims that bidirectional regularization, pulling model updates toward both old and new task knowledge, improves the stability-plasticity balance in continual learning.

desk verdict A useful two-phase plug-in for continual learning with consistent if modest gains, though undisclosed per-method FL hyperparameters and an overclaimed theoretical section keep it from being a clean accept. read the letter →

arxiv 2506.00477 v1 pith:OJYV5X6D submitted 2025-05-31 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords continuallearningcatastrophicforgettingstability-plasticitytrade-offbidirectionalregularizationknowledgedistillationmemoryreplayparameterdynamicarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Flashback Learning is a two-phase plugin for continual learning. In Phase 1 it quickly trains the old model on new data, saving that fast learner as 'plastic knowledge'. In Phase 2 it reinitializes to the old model and trains again while regularizing toward both the old model ('stable knowledge') and the Phase-1 learner. The paper argues this bidirectional pull makes the gradient target an interpolation between old and new responses, improving the stability-plasticity trade-off. Across replay, distillation, regularization, and architecture-expansion baselines, adding FL raises average accuracy by up to 4.91% in class-incremental and 3.51% in task-incremental settings under the same training budget.

What carries the argument

The mechanism is a pair of knowledge bases: SKB stores whatever stability information the host method already keeps (old model, memory logits, Fisher matrix, or frozen feature extractor), and PKB stores the corresponding information extracted from the Phase-1 primary model. The load-bearing identity is the gradient interpolation shown in Theorems 1-4, where the stability and plasticity losses combine into one term, e.g. for distillation: $\nabla_\theta L_{\mathrm{FL}}(\theta)=\nabla_\theta L_c(\theta)+(\alpha_s+\alpha_p)\nabla_\theta f(x;\theta)^\top\bigl(f(x;\theta)-\tfrac{\alpha_s f(x;\theta_s)+\alpha_p f(x;\theta_p)}{\alpha_s+\alpha_p}\bigr)$. This single term replaces the pure stability gradient with a pull toward an interpolation of old and new knowledge, which is what the paper identifies as the source of the improved trade-off.

What would settle it

Run FL on Split CIFAR-100 class-incremental with E1=10, E2=50 and compare against an ablation where the PKB is replaced by the old model's own responses, so the plasticity loss is identical to the stability loss; if the accuracy gap vanishes, the Phase-1 learner is not adding anything beyond retraining from the old checkpoint.

Watch

Extended reading notes

Core claim

The paper's central claim is that the forgetting problem can be reduced by giving the model an explicit learning target from the new task as well as the old one. FL does this in two phases: Phase 1 trains a 'primary' model on new data for few epochs; Phase 2 discards that model's weights but keeps its outputs or parameters as a plasticity target, reinitializes to the old model, and trains with a loss that pulls the network toward both targets. Theorems 1-4 show for each CL category that the FL gradient is the task gradient plus a single interpolation term, e.g. Eq. (25) for distillation, where the model is pulled toward a weighted average of stable and primary responses. The paper claims this gradient interpolation yields a better stability-plasticity balance than unidirectional regularization, and demonstrates lower stability-plasticity ratio, reduced forgetting, and accuracy gains on CIFAR and ImageNet benchmarks.

Load-bearing premise

The whole benefit rests on Phase 1 producing a genuinely better plasticity target than the old model within the small number of epochs E1; if Phase 1 is too short to learn anything useful or too long it starts to forget, the bidirectional pull degrades, and Table 6 shows E1=50 can hurt relative to E1=10.

Editorial extensions

If this is right

  • If FL works as claimed, any continual learner that already keeps stable knowledge can get a plasticity counterpart for roughly the same training cost.
  • Accuracy gains of up to about 4.9 points in class-incremental and 3.5 points in task-incremental settings should transfer to other benchmarks with similar task structure.
  • FL should lower forgetting and improve backward transfer for replay and distillation methods, not just final accuracy.
  • FL should remain beneficial under smaller replay buffers and across backbone changes from ResNet to vision transformers.
  • Plugging FL into architecture-expansion methods such as FOSTER and BEEF should improve their ImageNet-100 accuracy without requiring extra training epochs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A practical reading is that FL's balance knob is the ratio $\alpha_p/\alpha_s$, so the method could be exposed as a single tunable dial; the paper does not propose an automatic schedule for it.
  • Part of the gain could be a retraining effect, since Phase 2 restarts from the old weights and sees the same task data twice; isolating the pure bidirectional-loss contribution would require an ablation with the same two-pass schedule but no plasticity term.
  • Because PKB mirrors SKB, FL should extend to prompt-based or low-rank continual learners by treating the prompt or LoRA state as the knowledge base, a direction the paper lists as future work but does not test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper introduces Flashback Learning (FL), a two-phase plugin for continual learning. In Phase 1, FL starts from the model trained on previous tasks, trains briefly on the new task data to obtain a "primary model," and stores this as plastic knowledge. In Phase 2, FL reinitializes to the old model and trains again on the new task with a bidirectional regularization: the host method's stability loss plus a new plasticity loss that pulls the model toward the primary model. The paper derives gradient decompositions for four CL families (distillation, replay, parameter regularization, dynamic architecture) and claims that the resulting interpolation term enhances the stability-plasticity trade-off. Empirically, it reports average accuracy improvements for representative baselines in each family, for SOTA methods BEEF and FOSTER, and compares with two other two-phase plugin methods, with additional ablations on epochs, plasticity loss scale, memory size, and backbone architecture.

Significance. FL is a simple and modular idea with potentially broad applicability across CL method families. The paper has clear strengths: it evaluates at least one baseline from each of four CL categories, uses publicly available baseline codebases, transparently explains the X-DER Split-CIFAR-100 discrepancy, reports statistical significance tests, and includes extensive ablations. The gradient decompositions in Theorems 1-4 are algebraically correct as identities (up to typographical slips in the appendix). However, the headline empirical claim is weakened by the fact that FL-specific hyperparameters are adjusted empirically and not disclosed for the main tables, and the ablations show sensitivity of several accuracy points. The theoretical 'enhancement' claim is also only an interpretation of the interpolation form rather than a proven consequence. If a fixed, fully disclosed FL hyperparameter protocol reproduces the Table 1-2 improvements, this would be a useful and publishable plug-in method; the current manuscript does not yet establish that robustness.

major comments (2)
  1. [§6.7, Tables 6-7] The FL-specific hyperparameters E1, E2, and alpha_p are said to be 'adjusted empirically' but are never listed for the rows of Tables 1, 2, 4, or 5. This is load-bearing because the ablations show large sensitivity: Table 7 reports iCaRL on Split CIFAR-10 ranging from 73.52 (alpha_p=0.001) to 69.19 (alpha_p=1) against a 70.21 baseline, and Table 6 reports LUCIR dropping from 80.91 to 69.15 when E1/E2 are changed. Several headline gains are comparable to or smaller than this sensitivity, e.g., LUCIR+FL +1.02 on Split CIFAR-10 and X-DER+FL +0.45 on Split CIFAR-100. Without a full hyperparameter table and a fixed selection rule, one cannot exclude per-method, per-dataset selection of FL hyperparameters as the source of the reported improvements.
  2. [§4, after Eq. (25)] Theorems 1-4 are algebraic rewritings of the gradient of L_FL, and the statement that the interpolation term 'will enhance the stability-plasticity trade-off' does not follow from the displayed identity. The identity holds for any alpha_s, alpha_p, and any primary model, including one that has barely trained or one that has overfit the new task, so it cannot by itself predict changes in accuracy or forgetting. To make the theoretical contribution load-bearing, the authors should either present this as motivation (the FL gradient targets a convex combination of stable and plastic responses) or add explicit conditions under which the interpolation provably reduces forgetting or improves new-task accuracy.
minor comments (6)
  1. [§6.7.1-6.7.4] The headings 'Ablatin study' contain a typo and should read 'Ablation study'.
  2. [Theorem 4] Theorem 4 refers to the stable feature extractor as Eq. (6), but the stable feature extractor for dynamic architecture methods is defined in Eq. (8); Eq. (6) is the parameter-regularization stable knowledge.
  3. [Appendix C.5, Eq. (C.25)] Eq. (C.25) drops the parameter multipliers in the last two terms: 'eta alpha_s F_s + eta alpha_p F_p' should be 'eta (alpha_s F_s theta_s + alpha_p F_p theta_p)'. The subsequent derivation in Eq. (C.26) and Eq. (C.27) is correct, so this is a typographical error rather than a substantive flaw.
  4. [Tables 1, 2, 4] Tables 1, 2, and 4 do not report standard deviations or the number of seeds, while Table 5 does. Reporting error bars for the small gains (e.g., +0.10, +0.45) would help the reader assess whether the differences are within run-to-run variability.
  5. [Table 4 and §6.5.2] Table 4 mixes values directly reported from [24] with the authors' own runs; the text notes a replication baseline of 68.96 for FOSTER on ImageNet-100-B0-10, but the table itself does not mark which cells are own runs. Please mark own runs explicitly and provide the FL hyperparameters used for each SOTA integration.
  6. [Eq. (37), §6.8.2] The Stability-Plasticity Ratio is defined as Forgetting on Old Classes divided by Accuracy on New Classes, but the numerator uses Eq. (36), which averages forgetting over all previous tasks. Please clarify the wording so that the metric name matches the quantity computed.

Circularity Check

1 steps flagged · score 2.0 of 10

Theoretical 'enhancement' claim is a restatement of the loss definition; empirical gains are against external baselines, so circularity is minor.

  1. self definitional [Section 4.1, Eq. (25) (and analogous Eqs. 27, 29, 33)]
    "The interpolation term in gradient (25) drives the output f(x;θ) toward an interpolation between the stable response f(x;θ_s) and the primary new response f(x;θ_p). ... Therefore, the interpolation term guides the model output toward a balance between f(x;θ_s) and f(x;θ_p) and will enhance the stability-plasticity trade-off."

    Eq. (25) is obtained purely by expanding the gradient of L_FL = L_c + α_s L_s + α_p L_p, where L_s and L_p are defined as squared distances to f(x;θ_s) and f(x;θ_p), and then grouping the two linear terms. The 'interpolation' target (α_s f_s + α_p f_p)/(α_s+α_p) is exactly the minimizer of the combined regularization, so saying the gradient drives toward this interpolation is a restatement of the loss definition, not an independent derivation of improved stability-plasticity balance. The same algebraic rewriting occurs in Theorems 2-4. The empirical gains, however, are measured against external baselines and are not forced by this identity.

full rationale

The paper's derivation chain is largely an algebraic identity: given the FL loss in Eq. (19), the grouped gradient in Eq. (25) follows by expanding the two squared-error losses. Calling the grouped term 'Gradient Interpolation' and asserting it 'will enhance the stability-plasticity trade-off' is a definition-level restatement rather than a derived prediction; the actual trade-off benefit is an empirical claim. This is the only circularity-like element, and it does not contaminate the experimental comparisons, which use external baselines from mammoth and BEEF/FOSTER codebases. The self-citation to [20] is motivational and not load-bearing. The undisclosed per-method FL hyperparameters (Section 6.7) are a reproducibility/tuning concern, not a circularity of the derivation, since the reported gains are compared against baselines under the same training budget. Overall circularity is minor.

Assumptions & free parameters 2 free parameters · 4 assumptions · 2 invented entities

The central mechanism (Phase 1 plastic knowledge and Phase 2 bidirectional loss) rests on the empirical assumption that a short phase on new data yields useful plasticity targets; the paper provides ablations for this, not a proof. The theoretical analysis is an algebraic decomposition of the loss the authors define, not an independent derivation.

free parameters (2)
  • alpha_p (plasticity loss scaler) = 0.001 for iCaRL, 0.01 for LUCIR on Split CIFAR-10 (Table 7); not fixed across experiments
    Tuned per method and benchmark; accuracy varies from 69.19 to 73.52 as alpha_p ranges over 0 to 1 (Table 7).
  • E1 (Phase 1 epochs) = 10 used in most experiments; ablation tests 10, 50, and other splits (Table 6)
    Chosen so that E1+E2 equals the host training epochs; ablation shows E1=50 can reduce gains, so the value materially affects results.
assumptions (4)
  • domain assumption Stability and plasticity losses can be combined additively with the task loss while maintaining convergence.
    Section 3.3.2 defines L2 = Lc + alpha_s Ls + alpha_p Lp without a convergence guarantee; stability of the combined objective is assumed.
  • ad hoc to paper A short Phase 1 trained only for a few epochs on new data yields a primary model that is a useful plasticity target.
    Section 3.3.1 and the ablation in Table 6 show E1=10 works while E1=50 can hurt; this empirical assumption is not proven.
  • standard math Fisher information matrices F_s and F_p are treated as constant during the recursive SGD analysis in Appendix C.5.
    The closed-form trajectory (Eq. 30 and 31) assumes fixed FIMs, an approximation used to derive the steady-state interpolation argument.
  • standard math Standard gradient calculus and linear algebra, including the softmax gradient identity used in Theorem 4.
    Proofs in Appendix C rely on chain rule and the KL-divergence gradient for softmax outputs (e.g., Eq. C.17).
invented entities (2)
  • Plastic Knowledge Base (PKB)
    purpose: Stores the Phase-1 primary model's outputs, parameters, FIM, or module, as the plasticity target for Phase 2 regularization.
    The PKB is defined by the authors (Section 3.2); its usefulness is only demonstrated within this paper's experiments, with no external validation.
  • Primary model f(theta_p)
    purpose: A temporary model trained on new task data in Phase 1; serves as the source of plastic knowledge.
    Introduced in Section 3; its behavior is studied only in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Flashbacks to Harmonize Stability and Plasticity in Continual Learning." pith.science (2026). https://pith.science/paper/OJYV5X6D

@misc{pith2026250600477,
  author       = {Pith},
  title        = {Pith review of: Flashbacks to Harmonize Stability and Plasticity in Continual Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OJYV5X6D}},
  note         = {Machine review of arXiv:2506.00477}
}
read the original abstract

We introduce Flashback Learning (FL), a novel method designed to harmonize the stability and plasticity of models in Continual Learning (CL). Unlike prior approaches that primarily focus on regularizing model updates to preserve old information while learning new concepts, FL explicitly balances this trade-off through a bidirectional form of regularization. This approach effectively guides the model to swiftly incorporate new knowledge while actively retaining its old knowledge. FL operates through a two-phase training process and can be seamlessly integrated into various CL methods, including replay, parameter regularization, distillation, and dynamic architecture techniques. In designing FL, we use two distinct knowledge bases: one to enhance plasticity and another to improve stability. FL ensures a more balanced model by utilizing both knowledge bases to regularize model updates. Theoretically, we analyze how the FL mechanism enhances the stability-plasticity balance. Empirically, FL demonstrates tangible improvements over baseline methods within the same training budget. By integrating FL into at least one representative baseline from each CL category, we observed an average accuracy improvement of up to 4.91% in Class-Incremental and 3.51% in Task-Incremental settings on standard image classification benchmarks. Additionally, measurements of the stability-to-plasticity ratio confirm that FL effectively enhances this balance. FL also outperforms state-of-the-art CL methods on more challenging datasets like ImageNet.

Figures

Figures reproduced from arXiv: 2506.00477 by the authors.

Figure 1
Figure 1. Left: Former CL methods use existing knowledge from previous tasks to adjust model updates on new data. Right: FL method refines new knowledge out of new task data in Phase 1, then controls model updates by bidirectional regularization in Phase 2. the capacity to learn new data) [2]. The more effectively this balance is maintained, the more optimal the model’s performance in the continual learning setting. Towards t… view at source ↗
Figure 2
Figure 2. Flashback Learning overview; At task Tt , Phase 1- updates the old model on new data to obtain primary model f(·; θp ). Then, it extracts new knowledge from the primary model and stores it in PKB. Phase 2-flashbacks from primary to the old model f(·; θ ∗ t−1 ); while using stable knowledge and plastic knowledge to regularize model updates in a bidirectional flow and obtain new model f(·; θ ∗ t ). parameters subscrip… view at source ↗
Figure 3
Figure 3. Distillation methods keep f(·; θ ∗ t−1 ) at end of task t − 1 as stable knowledge for task Tt . 5 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Replay methods keep a small set of old samples [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Parameter regularization methods keep model parameters [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Dynamic architecture methods keep h(·; ϕ ∗ t−1 ) at the end of task t − 1 as stable knowledge for task Tt . 3.2. Plastic Knowledge Base Having defined stable knowledge per category in § 3.1, we explain how we extract plastic knowledge from new data in task Tt . As ment…
Figure 7
Figure 7. Figure 7: Phase 1 concludes by keeping f(·; θt ) as plastic knowledge for task Tt . 3.2.2. Memory Replay Compatible with old logits os or feature embeddings zs kept in replay method SKB (2), we evaluate memory samples x ∈ S by the primary model to obtain new logits op = f(x; θp …
Figure 8
Figure 8. Figure 8: Phase 1 evaluates memory samples by primary model to keep op or zp as plastic knowledge for task Tt . 3.2.3. Parameter Regularization To keep consistency with stable knowledge (6) defined in this category, we require model parameters and their FIM that have captured ta…
Figure 9
Figure 9. Figure 9: Phase 1 concludes by calculating Ft and taking {θt , Ft} as plastic knowledge for task Tt . 3.2.4. Dynamic Architecture Considering expansion stage (7) in architecture-based methods, new module m(·; φt ) attempts to acquire task Tt’s representation in the presence of s…
Figure 10
Figure 10. Figure 10: Phase 1 concludes by keeping the expanded module of m(·; φt ) as plastic knowledge for task Tt . 3.3. Flashback Learning Phases With a clear definition of stable and plastic knowledge in § 3.1 and § 3.2, we explain how SKB and PKB are incorporated in FL phases to harm…
Figure 11
Figure 11. Figure 11: Average accuracy (% ↑) reported on Split CIFAR-100 benchmark for host CL baseline EWC [28] (from parameter regularization category) and its improvement by P&C[5], AFAC [69], CAF [70], and FL mechanism. backbone and then applies distillation to prevent redundancy in fe…
Figure 12
Figure 12. Figure 12: Average Forgetting (↓), Backward Transfer (↑) and Forward Transfer (↑) are reported on Split-CIFAR-100. balance, we define Stability-Plasticity Ratio (SPR) as: \label {eq:spr} \text {SPR} = \frac {\text {Forgetting on Old Classes} \,}{\text {Accuracy on New Classes} \…
Figure 13
Figure 13. Figure 13: CKA Analysis on old and new test sets for oEWC and iCaRL in CL and FL settings [PITH_FULL_IMAGE:figures/full_fig_p029_13.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

94 extracted references · 79 canonical work pages

  1. [1]

    Robins, Catastrophic forgetting, rehearsal and pseudorehearsal, Connection Science 7 (2) (1995) 123–146

    A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal, Connection Science 7 (2) (1995) 123–146. 1

  2. [2]

    De Lange, R

    M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, T. Tuytelaars, A continual learning survey: Defying forgetting in classification tasks, IEEE Conf. Comput. Vis. Pattern Recog. 44 (7) (2021) 3366–3385. 2

  3. [3]

    Chaudhry, P

    A. Chaudhry, P. K. Dokania, T. Ajanthan, P. H. Torr, Riemannian walk for incremental learning: Understanding forgetting and intransigence, in: Eur. Conf. Comput. Vis., 2018, pp. 532–547. 2, 17, 20

  4. [4]

    X. Nie, S. Xu, X. Liu, G. Meng, C. Huo, S. Xiang, Bilateral memory consolidation for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 16026–16035. 2

  5. [5]

    Schwarz, W

    J. Schwarz, W. Czarnecki, J. Luketina, A. Grabska-Barwinska, Y . W. Teh, R. Pascanu, R. Hadsell, Progress & compress: A scalable framework for continual learning, in: Int. Conf. Mach. Learn., 2018, pp. 4528–4537. 2, 3, 6, 17, 18, 21, 28

  6. [6]

    Z. Li, D. Hoiem, Learning without Forgetting, IEEE Trans. Pattern Anal. Mach. Intell. (2017). 2, 3, 5, 16, 20

  7. [7]

    Douillard, Y

    A. Douillard, Y . Chen, A. Dapogny, M. Cord, Plop: Learning without forgetting for continual semantic segmen- tation, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 4040–4050. 2

  8. [8]

    K. Roy, C. Simon, P. Moghadam, M. Harandi, Subspace distillation for continual learning, Neural Networks (2023). 2, 16

Show all 94 references
  1. [9]

    T. L. Hayes, G. P. Krishnan, M. Bazhenov, H. T. Siegelmann, T. J. Sejnowski, C. Kanan, Replay in deep learning: Current approaches and missing biological elements, Neural computation 33 (11) (2021) 2908–2950. 2

  2. [10]

    Bonicelli, M

    L. Bonicelli, M. Boschini, A. Porrello, C. Spampinato, S. Calderara, On the effectiveness of lipschitz-driven rehearsal in continual learning, Adv. Neural Inform. Process. Syst. 35 (2022) 31886–31901. 2

  3. [11]

    Rebuffi, A

    S.-A. Rebuffi, A. Kolesnikov, G. Sperl, C. H. Lampert, iCaRL: Incremental Classifier and Representation Learn- ing, in: IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2001–2010. 2, 3, 16, 19, 20, 21, 22, 23, 24, 25, 26, 28

  4. [12]

    H. Shin, J. K. Lee, J. Kim, J. Kim, Continual Learning with Deep Generative Replay, in: Adv. Neural Inform. Process. Syst., 2017, pp. 1–10. 2

  5. [13]

    G. M. Van de Ven, H. T. Siegelmann, A. S. Tolias, Brain-inspired replay for continual learning with artificial neural networks, Nature Communications 11 (1) (2020) 4069. 2, 16

  6. [14]

    Mundt, I

    M. Mundt, I. Pliushch, S. Majumder, Y . Hong, V . Ramesh, Unified probabilistic deep continual learning through generative replay and open set recognition, Journal of Imaging 8 (4) (2022) 93. 2

  7. [15]

    D. Kim, B. Han, On the stability-plasticity dilemma of class-incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 20196–20204. 2, 16

  8. [16]

    G. Yang, F. Pan, W.-B. Gan, Stably maintained dendritic spines are associated with lifelong memories, Nature 462 (7275) (2009) 920–924. 2

  9. [17]

    D. Ji, M. A. Wilson, Coordinated memory replay in the visual cortex and hippocampus during sleep, Nature neuroscience 10 (1) (2007) 100–107. 2

  10. [18]

    O. C. Gonzalez, Y . Sokolov, G. P. Krishnan, J. E. Delanois, M. Bazhenov, Can sleep protect memories from catastrophic forgetting?, eLife 9 (2020) e51005. 2

  11. [19]

    D. J. Bridge, J. L. V oss, Hippocampal binding of novel information with dominant memory traces can support both memory stability and change, Journal of Neuroscience 34 (6) (2014) 2203–2213. 2 36

  12. [20]

    Mahmoodi, M

    L. Mahmoodi, M. Harandi, P. Moghadam, Flashback for continual learning, in: Proceedings of Int. Conf. Com- put. Vis., 2023, pp. 3434–3443. 3

  13. [21]

    S. Hou, X. Pan, C. C. Loy, Z. Wang, D. Lin, Learning a Unified Classifier Incrementally via Rebalancing, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 831–839. 3, 5, 10, 16, 19, 20, 21, 22, 24, 25, 26, 28, 30

  14. [22]

    Boschini, L

    M. Boschini, L. Bonicelli, P. Buzzega, A. Porrello, S. Calderara, Class-incremental continual learning into the extended der-verse, IEEE Trans. Pattern Anal. Mach. Intell. 45 (5) (2022) 5497–5512. 3, 16, 19, 21, 22, 24

  15. [23]

    Wang, D.-W

    F.-Y . Wang, D.-W. Zhou, H.-J. Ye, D.-C. Zhan, Foster: Feature boosting and compression for class-incremental learning, in: Eur. Conf. Comput. Vis., Springer, 2022, pp. 398–414. 3, 7, 10, 17, 19, 20, 22, 23

  16. [24]

    Wang, D.-W

    F.-Y . Wang, D.-W. Zhou, L. Liu, H.-J. Ye, Y . Bian, D.-C. Zhan, P. Zhao, Beef: Bi-compatible class-incremental learning via energy-based expansion and fusion, in: Int. Conf. Learn. Represent., 2023, pp. 1–25. 3, 18, 19, 20, 22, 23

  17. [25]

    Douillard, M

    A. Douillard, M. Cord, C. Ollion, T. Robert, E. Valle, PODNet: Pooled outputs distillation for small-tasks incremental learning, in: Eur. Conf. Comput. Vis., 2020, pp. 86–102. 5, 16, 19, 23

  18. [26]

    Buzzega, M

    P. Buzzega, M. Boschini, A. Porrello, D. Abati, S. Calderara, Dark experience for general continual learning: a strong, simple baseline, Adv. Neural Inform. Process. Syst. 33 (2020) 15920–15930. 6, 10, 16, 19, 21, 22, 26

  19. [27]

    Iscen, J

    A. Iscen, J. Zhang, S. Lazebnik, C. Schmid, Memory-efficient incremental learning through feature adaptation, in: Eur. Conf. Comput. Vis., 2020, pp. 699–715. 6, 16

  20. [28]

    Kirkpatrick, R

    J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ra- malho, A. Grabska-Barwinska, et al., Overcoming catastrophic forgetting in neural networks, Proceedings of the national academy of sciences 114 (13) (2017) 3521–3526. 6, ...

  21. [29]

    S. Yan, J. Xie, X. He, DER: Dynamically Expandable Representation for Class Incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 3014–3023. 7, 17, 23

  22. [30]

    Mundt, Y

    M. Mundt, Y . Hong, I. Pliushch, V . Ramesh, A wholistic view of continual learning with deep neural networks: Forgotten lessons and the bridge to active and open world learning, Neural Networks 160 (2023) 306–336. 15

  23. [31]

    K. Roy, C. Simon, P. Moghadam, M. Harandi, Cl3: Generalization of contrastive loss for lifelong learning, Journal of Imaging 9 (12) (2023). 16

  24. [32]

    K. Roy, P. Moghadam, M. Harandi, L3dmc: Lifelong learning using distillation via mixed-curvature space, in: Int. Conf. on Med. Image Computing and Computer-Assisted Intervention, 2023, pp. 123–133. 16

  25. [33]

    Simon, P

    C. Simon, P. Koniusz, M. Harandi, On Learning the Geodesic Path for Incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 1591–1600. 16, 19

  26. [34]

    M. Kang, J. Park, B. Han, Class-incremental learning by knowledge distillation with adaptive feature consolida- tion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16071–16080. 16, 19

  27. [35]

    Zhu, X.-Y

    F. Zhu, X.-Y . Zhang, C. Wang, F. Yin, C.-L. Liu, Prototype augmentation and self-supervision for incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 5871–5880. 16

  28. [36]

    Knights, P

    J. Knights, P. Moghadam, M. Ramezani, S. Sridharan, C. Fookes, Incloud: Incremental learning for point cloud place recognition, in: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2022, pp. 8559–8566. 16

  29. [37]

    H. Yin, P. Molchanov, J. M. Alvarez, Z. Li, A. Mallya, D. Hoiem, N. K. Jha, J. Kautz, Dreaming to distill: Data- free knowledge transfer via deepinversion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 8715–8724. 16 37

  30. [38]

    X. Li, S. Wang, J. Sun, Z. Xu, Memory efficient data-free distillation for continual learning, Pattern Recognition 144 (2023) 109875. 16

  31. [39]

    E. Fini, S. Lathuiliere, E. Sangineto, M. Nabi, E. Ricci, Online continual learning under extreme memory con- straints, in: Eur. Conf. Comput. Vis., Springer, 2020, pp. 720–735. 16

  32. [40]

    E. Fini, V . G. T. Da Costa, X. Alameda-Pineda, E. Ricci, K. Alahari, J. Mairal, Self-supervised models are continual learners, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 9621–9630. 16

  33. [41]

    Gomez-Villa, B

    A. Gomez-Villa, B. Twardowski, L. Yu, A. D. Bagdanov, J. van de Weijer, Continually learning self-supervised representations with projected functional regularization, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3867–3877. 16

  34. [42]

    Q. Gu, D. Shim, F. Shkurti, Preserving linear separability in continual learning by backward feature projection, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 24286–24295. 16, 19

  35. [43]

    Riemer, I

    M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y . Tu, , G. Tesauro, Learning to learn without forgetting by maximizing transfer and minimizing interference, in: Int. Conf. Learn. Represent., 2019, pp. 1–31. 16, 23

  36. [44]

    Caccia, R

    L. Caccia, R. Aljundi, N. Asadi, T. Tuytelaars, J. Pineau, E. Belilovsky, New insights on reducing abrupt repre- sentation change in online continual learning, in: Int. Conf. Learn. Represent., 2022, pp. 1–27. 16

  37. [45]

    J. Bang, H. Kim, Y . Yoo, J.-W. Ha, J. Choi, Rainbow memory: Continual learning with a memory of diverse samples, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 8218–8227. 16

  38. [46]

    Y . Liu, Y . Su, A.-A. Liu, B. Schiele, Q. Sun, Mnemonics Training: Multi-Class Incremental Learning without Forgetting, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 12245–12254. 16

  39. [47]

    S. Ho, M. Liu, L. Du, L. Gao, Y . Xiang, Prototype-guided memory replay for continual learning, IEEE Trans. Neural Net. Learn. Sys. (2023). 16

  40. [48]

    Aljundi, K

    R. Aljundi, K. Kelchtermans, T. Tuytelaars, Task-free continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 11254–11263. 16

  41. [49]

    Aljundi, E

    R. Aljundi, E. Belilovsky, T. Tuytelaars, L. Charlin, M. Caccia, M. Lin, L. Page-Caccia, Online continual learn- ing with maximal interfered retrieval, Adv. Neural Inform. Process. Syst. 32 (2019). 16

  42. [50]

    Caccia, E

    L. Caccia, E. Belilovsky, M. Caccia, J. Pineau, Online learned continual compression with adaptive quantization modules, in: Int. Conf. Mach. Learn., PMLR, 2020, pp. 1240–1250. 16

  43. [51]

    Y . Oh, D. Baek, B. Ham, Alife: Adaptive logit regularizer and feature replay for incremental semantic segmen- tation, Adv. Neural Inform. Process. Syst. 35 (2022) 14516–14528. 16

  44. [52]

    L. Wang, X. Zhang, K. Yang, L. Yu, C. Li, L. HONG, S. Zhang, Z. Li, Y . Zhong, J. Zhu, Memory replay with data compression for continual learning, in: Int. Conf. Learn. Represent., 2022, pp. 1–25. 16

  45. [53]

    X. Liu, C. Wu, M. Menta, L. Herranz, B. Raducanu, A. D. Bagdanov, S. Jui, J. v. de Weijer, Generative feature replay for class-incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 226–227. 16

  46. [54]

    F. Ye, A. G. Bors, Learning latent representations across multiple data domains using lifelong vaegan, in: Eur. Conf. Comput. Vis., 2020, pp. 777–795. 16

  47. [55]

    Binici, S

    K. Binici, S. Aggarwal, N. T. Pham, K. Leman, T. Mitra, Robust and resource-efficient data-free knowledge distillation by generative pseudo replay, in: AAAI, V ol. 36, 2022, pp. 6089–6096. 16

  48. [56]

    G. M. Van De Ven, Z. Li, A. S. Tolias, Class-incremental learning with generative classifiers, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 3611–3620. 16 38

  49. [57]

    Y . Cong, M. Zhao, J. Li, S. Wang, L. Carin, Gan memory with no forgetting, Adv. Neural Inform. Process. Syst. 33 (2020) 16481–16494. 16

  50. [58]

    T. Li, G. Pang, X. Bai, J. Zheng, L. Zhou, X. Ning, Learning adversarial semantic embeddings for zero-shot recognition in open worlds, Pattern Recognition 149 (2024) 110258. 16

  51. [59]

    Lopez-Paz, M

    D. Lopez-Paz, M. Ranzato, Gradient Episodic Memory for Continual Learning, in: Adv. Neural Inform. Process. Syst., 2017, pp. 1–10. 16, 20

  52. [60]

    Chaudhry, M

    A. Chaudhry, M. Ranzato, M. Rohrbach, M. Elhoseiny, Efficient lifelong learning with a-GEM, in: Int. Conf. Learn. Represent., 2019, pp. 1–20. 16

  53. [61]

    S. Jung, H. Ahn, S. Cha, T. Moon, Continual learning with node-importance based adaptive group sparse regu- larization, Adv. Neural Inform. Process. Syst. 33 (2020) 3647–3658. 17

  54. [62]

    J. Lee, H. G. Hong, D. Joo, J. Kim, Continual learning with extended kronecker-factored approximate curvature, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9001–9010. 17

  55. [63]

    Zenke, B

    F. Zenke, B. Poole, S. Ganguli, Continual Learning Through Synaptic Intelligence, in: Int. Conf. Mach. Learn., 2017, pp. 1–9. 17, 19

  56. [64]

    Aljundi, F

    R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, T. Tuytelaars, Memory Aware Synapses: Learning what (not) to forget, in: Eur. Conf. Comput. Vis., 2018, pp. 1–16. 17

  57. [65]

    H. Ran, W. Li, L. Li, S. Tian, X. Ning, P. Tiwari, Learning optimal inter-class margin adaptively for few- shot class-incremental learning via neural collapse-based meta-learning, Information Processing & Management 61 (3) (2024) 103664. 17

  58. [66]

    Pelosin, S

    F. Pelosin, S. Jha, A. Torsello, B. Raducanu, J. van de Weijer, Towards exemplar-free continual learning in vision transformers: an account of attention, functional and weight regularization, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3820–3829. 17

  59. [67]

    Soutif-Cormerais, A

    A. Soutif-Cormerais, A. Carta, J. van de Weijer, Improving online continual learning performance and stability with temporal ensembles, in: S. Chandar, R. Pascanu, H. Sedghi, D. Precup (Eds.), Proceedings of Machine Learn. Research, V ol. 232, PMLR, 2023, pp. 828–845. 17

  60. [68]

    Gupta, K

    G. Gupta, K. Yadav, L. Paull, La-maml: Look-ahead meta learning for continual learning. 2020, URL https://arxiv. org/abs (2007). 17

  61. [69]

    L. Wang, M. Zhang, Z. Jia, Q. Li, C. Bao, K. Ma, J. Zhu, Y . Zhong, Afec: Active forgetting of negative transfer in continual learning, Adv. Neural Inform. Process. Syst. 34 (2021) 22379–22391. 18

  62. [70]

    L. Wang, X. Zhang, Q. Li, M. Zhang, H. Su, J. Zhu, Y . Zhong, Incorporating neuro-inspired adaptability for continual learning in artificial intelligence, Nature Machine Intelligence 5 (12) (2023) 1356–1368. 18

  63. [71]

    Serra, D

    J. Serra, D. Suris, M. Miron, A. Karatzoglou, Overcoming catastrophic forgetting with hard attention to the task, in: Int. Conf. Learn. Represent., PMLR, 2018, pp. 4548–4557. 17

  64. [72]

    Ostapenko, M

    O. Ostapenko, M. Puscas, T. Klein, P. Jahnichen, M. Nabi, Learning to Remember: A Synaptic Plasticity Driven Framework for Continual Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 11313–11321. 17

  65. [73]

    J. Yoon, E. Yang, J. Lee, S. J. Hwang, Lifelong Learning with Dynamically Expandable Networks, in: Int. Conf. Learn. Represent., 2018, pp. 1–11. 17

  66. [74]

    X. Li, Y . Zhou, T. Wu, R. Socher, C. Xiong, Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting, in: Int. Conf. Mach. Learn., PMLR, 2019, pp. 3925–3934. 17 39

  67. [75]

    G. Lin, H. Chu, H. Lai, Towards better plasticity-stability trade-off in incremental learning: A simple linear connector, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 89–98. 18

  68. [76]

    Krizhevsky, G

    A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, https://www.cs.toronto.edu/ kriz/learning-features-2009-TR.pdf (2009) 3–58. 19

  69. [77]

    J. Wu, Q. Zhang, G. Xu, Tiny imagenet challenge, Technical Report (2017). 19

  70. [78]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: IEEE Conf. Comput. Vis. Pattern Recog., Ieee, 2009, pp. 248–255. 19

  71. [79]

    Y . Wu, Y . Chen, L. Wang, Y . Ye, Z. Liu, Y . Guo, Y . Fu, Large scale incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 374–382. 19, 22, 23, 26

  72. [80]

    S. Kim, L. Noci, A. Orvieto, T. Hofmann, Achieving a better stability-plasticity trade-off via auxiliary networks in continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 11930–11939. 20, 24

  73. [81]

    Y . Liu, B. Schiele, Q. Sun, Adaptive aggregation networks for class-incremental learning, in: IEEE Conf. Com- put. Vis. Pattern Recog., 2021, pp. 2544–2553. 20, 24

  74. [82]

    K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 770–778. 22, 23

  75. [83]

    F. M. Castro, M. J. Marín-Jiménez, N. Guil, C. Schmid, K. Alahari, End-to-End Incremental Learning, in: Eur. Conf. Comput. Vis., 2018, pp. 1–16. 22, 26

  76. [84]

    H. Ahn, J. Kwak, S. Lim, H. Bang, H. Kim, T. Moon, SS-IL: Separated Softmax for Incremental Learning, in: Int. Conf. Comput. Vis., 2021, pp. 844–853. 22

  77. [85]

    B. Zhao, X. Xiao, G. Gan, B. Zhang, S.-T. Xia, Maintaining Discrimination and Fairness in Class incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 13208–13207. 23

  78. [86]

    Douillard, A

    A. Douillard, A. Ramé, G. Couairon, M. Cord, Dytox: Transformers for continual learning with dynamic token expansion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 9285–9295. 23

  79. [87]

    Dosovitskiy, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)

    A. Dosovitskiy, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020). 26

  80. [88]

    Kornblith, M

    S. Kornblith, M. Norouzi, H. Lee, G. Hinton, Similarity of neural network representations revisited, in: Int. Conf. Mach. Learn., PMLR, 2019, pp. 3519–3529. 27

  81. [89]

    Yosinski, J

    J. Yosinski, J. Clune, Y . Bengio, H. Lipson, How transferable are features in deep neural networks?, Adv. Neural Inform. Process. Syst. 27 (2014). 28

  82. [90]

    Z. Wang, Z. Zhang, C.-Y . Lee, H. Zhang, R. Sun, X. Ren, G. Su, V . Perot, J. Dy, T. Pfister, Learning to prompt for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 139–149. 28

  83. [91]

    J. S. Smith, L. Karlinsky, V . Gutta, P. Cascante-Bonilla, D. Kim, A. Arbelle, R. Panda, R. Feris, Z. Kira, Coda- prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 11909–11919. 28

  84. [92]

    Liang, W.-J

    Y .-S. Liang, W.-J. Li, Inflora: Interference-free low-rank adaptation for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 23638–23647. 28

  85. [93]

    X. Wang, T. Chen, Q. Ge, H. Xia, R. Bao, R. Zheng, Q. Zhang, T. Gui, X.-J. Huang, Orthogonal subspace learning for language model continual learning, in: Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 10658–10671. 28

  86. [94]

    C. M. Bishop, H. Bishop, Deep learning: Foundations and concepts, Springer Nature, 2023. 33 40

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.