REVIEW 2 major objections 6 minor 94 references
Flashbacks to Harmonize Stability and Plasticity in Continual Learning
T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Flashback Learning claims that bidirectional regularization, pulling model updates toward both old and new task knowledge, improves the stability-plasticity balance in continual learning.
desk verdict A useful two-phase plug-in for continual learning with consistent if modest gains, though undisclosed per-method FL hyperparameters and an overclaimed theoretical section keep it from being a clean accept. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is a pair of knowledge bases: SKB stores whatever stability information the host method already keeps (old model, memory logits, Fisher matrix, or frozen feature extractor), and PKB stores the corresponding information extracted from the Phase-1 primary model. The load-bearing identity is the gradient interpolation shown in Theorems 1-4, where the stability and plasticity losses combine into one term, e.g. for distillation: $\nabla_\theta L_{\mathrm{FL}}(\theta)=\nabla_\theta L_c(\theta)+(\alpha_s+\alpha_p)\nabla_\theta f(x;\theta)^\top\bigl(f(x;\theta)-\tfrac{\alpha_s f(x;\theta_s)+\alpha_p f(x;\theta_p)}{\alpha_s+\alpha_p}\bigr)$. This single term replaces the pure stability gradient with a pull toward an interpolation of old and new knowledge, which is what the paper identifies as the source of the improved trade-off.
What would settle it
Run FL on Split CIFAR-100 class-incremental with E1=10, E2=50 and compare against an ablation where the PKB is replaced by the old model's own responses, so the plasticity loss is identical to the stability loss; if the accuracy gap vanishes, the Phase-1 learner is not adding anything beyond retraining from the old checkpoint.
Extended reading notes
Core claim
The paper's central claim is that the forgetting problem can be reduced by giving the model an explicit learning target from the new task as well as the old one. FL does this in two phases: Phase 1 trains a 'primary' model on new data for few epochs; Phase 2 discards that model's weights but keeps its outputs or parameters as a plasticity target, reinitializes to the old model, and trains with a loss that pulls the network toward both targets. Theorems 1-4 show for each CL category that the FL gradient is the task gradient plus a single interpolation term, e.g. Eq. (25) for distillation, where the model is pulled toward a weighted average of stable and primary responses. The paper claims this gradient interpolation yields a better stability-plasticity balance than unidirectional regularization, and demonstrates lower stability-plasticity ratio, reduced forgetting, and accuracy gains on CIFAR and ImageNet benchmarks.
Load-bearing premise
The whole benefit rests on Phase 1 producing a genuinely better plasticity target than the old model within the small number of epochs E1; if Phase 1 is too short to learn anything useful or too long it starts to forget, the bidirectional pull degrades, and Table 6 shows E1=50 can hurt relative to E1=10.
Editorial extensions
If this is right
- If FL works as claimed, any continual learner that already keeps stable knowledge can get a plasticity counterpart for roughly the same training cost.
- Accuracy gains of up to about 4.9 points in class-incremental and 3.5 points in task-incremental settings should transfer to other benchmarks with similar task structure.
- FL should lower forgetting and improve backward transfer for replay and distillation methods, not just final accuracy.
- FL should remain beneficial under smaller replay buffers and across backbone changes from ResNet to vision transformers.
- Plugging FL into architecture-expansion methods such as FOSTER and BEEF should improve their ImageNet-100 accuracy without requiring extra training epochs.
Reading between the lines
- A practical reading is that FL's balance knob is the ratio $\alpha_p/\alpha_s$, so the method could be exposed as a single tunable dial; the paper does not propose an automatic schedule for it.
- Part of the gain could be a retraining effect, since Phase 2 restarts from the old weights and sees the same task data twice; isolating the pure bidirectional-loss contribution would require an ablation with the same two-pass schedule but no plasticity term.
- Because PKB mirrors SKB, FL should extend to prompt-based or low-rank continual learners by treating the prompt or LoRA state as the knowledge base, a direction the paper lists as future work but does not test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Flashback Learning (FL), a two-phase plugin for continual learning. In Phase 1, FL starts from the model trained on previous tasks, trains briefly on the new task data to obtain a "primary model," and stores this as plastic knowledge. In Phase 2, FL reinitializes to the old model and trains again on the new task with a bidirectional regularization: the host method's stability loss plus a new plasticity loss that pulls the model toward the primary model. The paper derives gradient decompositions for four CL families (distillation, replay, parameter regularization, dynamic architecture) and claims that the resulting interpolation term enhances the stability-plasticity trade-off. Empirically, it reports average accuracy improvements for representative baselines in each family, for SOTA methods BEEF and FOSTER, and compares with two other two-phase plugin methods, with additional ablations on epochs, plasticity loss scale, memory size, and backbone architecture.
Significance. FL is a simple and modular idea with potentially broad applicability across CL method families. The paper has clear strengths: it evaluates at least one baseline from each of four CL categories, uses publicly available baseline codebases, transparently explains the X-DER Split-CIFAR-100 discrepancy, reports statistical significance tests, and includes extensive ablations. The gradient decompositions in Theorems 1-4 are algebraically correct as identities (up to typographical slips in the appendix). However, the headline empirical claim is weakened by the fact that FL-specific hyperparameters are adjusted empirically and not disclosed for the main tables, and the ablations show sensitivity of several accuracy points. The theoretical 'enhancement' claim is also only an interpretation of the interpolation form rather than a proven consequence. If a fixed, fully disclosed FL hyperparameter protocol reproduces the Table 1-2 improvements, this would be a useful and publishable plug-in method; the current manuscript does not yet establish that robustness.
major comments (2)
- [§6.7, Tables 6-7] The FL-specific hyperparameters E1, E2, and alpha_p are said to be 'adjusted empirically' but are never listed for the rows of Tables 1, 2, 4, or 5. This is load-bearing because the ablations show large sensitivity: Table 7 reports iCaRL on Split CIFAR-10 ranging from 73.52 (alpha_p=0.001) to 69.19 (alpha_p=1) against a 70.21 baseline, and Table 6 reports LUCIR dropping from 80.91 to 69.15 when E1/E2 are changed. Several headline gains are comparable to or smaller than this sensitivity, e.g., LUCIR+FL +1.02 on Split CIFAR-10 and X-DER+FL +0.45 on Split CIFAR-100. Without a full hyperparameter table and a fixed selection rule, one cannot exclude per-method, per-dataset selection of FL hyperparameters as the source of the reported improvements.
- [§4, after Eq. (25)] Theorems 1-4 are algebraic rewritings of the gradient of L_FL, and the statement that the interpolation term 'will enhance the stability-plasticity trade-off' does not follow from the displayed identity. The identity holds for any alpha_s, alpha_p, and any primary model, including one that has barely trained or one that has overfit the new task, so it cannot by itself predict changes in accuracy or forgetting. To make the theoretical contribution load-bearing, the authors should either present this as motivation (the FL gradient targets a convex combination of stable and plastic responses) or add explicit conditions under which the interpolation provably reduces forgetting or improves new-task accuracy.
minor comments (6)
- [§6.7.1-6.7.4] The headings 'Ablatin study' contain a typo and should read 'Ablation study'.
- [Theorem 4] Theorem 4 refers to the stable feature extractor as Eq. (6), but the stable feature extractor for dynamic architecture methods is defined in Eq. (8); Eq. (6) is the parameter-regularization stable knowledge.
- [Appendix C.5, Eq. (C.25)] Eq. (C.25) drops the parameter multipliers in the last two terms: 'eta alpha_s F_s + eta alpha_p F_p' should be 'eta (alpha_s F_s theta_s + alpha_p F_p theta_p)'. The subsequent derivation in Eq. (C.26) and Eq. (C.27) is correct, so this is a typographical error rather than a substantive flaw.
- [Tables 1, 2, 4] Tables 1, 2, and 4 do not report standard deviations or the number of seeds, while Table 5 does. Reporting error bars for the small gains (e.g., +0.10, +0.45) would help the reader assess whether the differences are within run-to-run variability.
- [Table 4 and §6.5.2] Table 4 mixes values directly reported from [24] with the authors' own runs; the text notes a replication baseline of 68.96 for FOSTER on ImageNet-100-B0-10, but the table itself does not mark which cells are own runs. Please mark own runs explicitly and provide the FL hyperparameters used for each SOTA integration.
- [Eq. (37), §6.8.2] The Stability-Plasticity Ratio is defined as Forgetting on Old Classes divided by Accuracy on New Classes, but the numerator uses Eq. (36), which averages forgetting over all previous tasks. Please clarify the wording so that the metric name matches the quantity computed.
Circularity Check
Theoretical 'enhancement' claim is a restatement of the loss definition; empirical gains are against external baselines, so circularity is minor.
-
self definitional
[Section 4.1, Eq. (25) (and analogous Eqs. 27, 29, 33)]
"The interpolation term in gradient (25) drives the output f(x;θ) toward an interpolation between the stable response f(x;θ_s) and the primary new response f(x;θ_p). ... Therefore, the interpolation term guides the model output toward a balance between f(x;θ_s) and f(x;θ_p) and will enhance the stability-plasticity trade-off."
Eq. (25) is obtained purely by expanding the gradient of L_FL = L_c + α_s L_s + α_p L_p, where L_s and L_p are defined as squared distances to f(x;θ_s) and f(x;θ_p), and then grouping the two linear terms. The 'interpolation' target (α_s f_s + α_p f_p)/(α_s+α_p) is exactly the minimizer of the combined regularization, so saying the gradient drives toward this interpolation is a restatement of the loss definition, not an independent derivation of improved stability-plasticity balance. The same algebraic rewriting occurs in Theorems 2-4. The empirical gains, however, are measured against external baselines and are not forced by this identity.
full rationale
The paper's derivation chain is largely an algebraic identity: given the FL loss in Eq. (19), the grouped gradient in Eq. (25) follows by expanding the two squared-error losses. Calling the grouped term 'Gradient Interpolation' and asserting it 'will enhance the stability-plasticity trade-off' is a definition-level restatement rather than a derived prediction; the actual trade-off benefit is an empirical claim. This is the only circularity-like element, and it does not contaminate the experimental comparisons, which use external baselines from mammoth and BEEF/FOSTER codebases. The self-citation to [20] is motivational and not load-bearing. The undisclosed per-method FL hyperparameters (Section 6.7) are a reproducibility/tuning concern, not a circularity of the derivation, since the reported gains are compared against baselines under the same training budget. Overall circularity is minor.
Assumptions & free parameters
free parameters (2)
- alpha_p (plasticity loss scaler) =
0.001 for iCaRL, 0.01 for LUCIR on Split CIFAR-10 (Table 7); not fixed across experiments
- E1 (Phase 1 epochs) =
10 used in most experiments; ablation tests 10, 50, and other splits (Table 6)
assumptions (4)
- domain assumption Stability and plasticity losses can be combined additively with the task loss while maintaining convergence.
- ad hoc to paper A short Phase 1 trained only for a few epochs on new data yields a primary model that is a useful plasticity target.
- standard math Fisher information matrices F_s and F_p are treated as constant during the recursive SGD analysis in Appendix C.5.
- standard math Standard gradient calculus and linear algebra, including the softmax gradient identity used in Theorem 4.
invented entities (2)
-
Plastic Knowledge Base (PKB)
-
Primary model f(theta_p)
Cite this review
Pith. "Pith review of Flashbacks to Harmonize Stability and Plasticity in Continual Learning." pith.science (2026). https://pith.science/paper/OJYV5X6D
@misc{pith2026250600477,
author = {Pith},
title = {Pith review of: Flashbacks to Harmonize Stability and Plasticity in Continual Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/OJYV5X6D}},
note = {Machine review of arXiv:2506.00477}
}
read the original abstract
We introduce Flashback Learning (FL), a novel method designed to harmonize the stability and plasticity of models in Continual Learning (CL). Unlike prior approaches that primarily focus on regularizing model updates to preserve old information while learning new concepts, FL explicitly balances this trade-off through a bidirectional form of regularization. This approach effectively guides the model to swiftly incorporate new knowledge while actively retaining its old knowledge. FL operates through a two-phase training process and can be seamlessly integrated into various CL methods, including replay, parameter regularization, distillation, and dynamic architecture techniques. In designing FL, we use two distinct knowledge bases: one to enhance plasticity and another to improve stability. FL ensures a more balanced model by utilizing both knowledge bases to regularize model updates. Theoretically, we analyze how the FL mechanism enhances the stability-plasticity balance. Empirically, FL demonstrates tangible improvements over baseline methods within the same training budget. By integrating FL into at least one representative baseline from each CL category, we observed an average accuracy improvement of up to 4.91% in Class-Incremental and 3.51% in Task-Incremental settings on standard image classification benchmarks. Additionally, measurements of the stability-to-plasticity ratio confirm that FL effectively enhances this balance. FL also outperforms state-of-the-art CL methods on more challenging datasets like ImageNet.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Robins, Catastrophic forgetting, rehearsal and pseudorehearsal, Connection Science 7 (2) (1995) 123–146
A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal, Connection Science 7 (2) (1995) 123–146. 1
1995
-
[2]
De Lange, R
M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, T. Tuytelaars, A continual learning survey: Defying forgetting in classification tasks, IEEE Conf. Comput. Vis. Pattern Recog. 44 (7) (2021) 3366–3385. 2
2021
-
[3]
Chaudhry, P
A. Chaudhry, P. K. Dokania, T. Ajanthan, P. H. Torr, Riemannian walk for incremental learning: Understanding forgetting and intransigence, in: Eur. Conf. Comput. Vis., 2018, pp. 532–547. 2, 17, 20
2018
-
[4]
X. Nie, S. Xu, X. Liu, G. Meng, C. Huo, S. Xiang, Bilateral memory consolidation for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 16026–16035. 2
2023
-
[5]
Schwarz, W
J. Schwarz, W. Czarnecki, J. Luketina, A. Grabska-Barwinska, Y . W. Teh, R. Pascanu, R. Hadsell, Progress & compress: A scalable framework for continual learning, in: Int. Conf. Mach. Learn., 2018, pp. 4528–4537. 2, 3, 6, 17, 18, 21, 28
2018
-
[6]
Z. Li, D. Hoiem, Learning without Forgetting, IEEE Trans. Pattern Anal. Mach. Intell. (2017). 2, 3, 5, 16, 20
2017
-
[7]
Douillard, Y
A. Douillard, Y . Chen, A. Dapogny, M. Cord, Plop: Learning without forgetting for continual semantic segmen- tation, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 4040–4050. 2
2021
-
[8]
K. Roy, C. Simon, P. Moghadam, M. Harandi, Subspace distillation for continual learning, Neural Networks (2023). 2, 16
2023
Show all 94 references
-
[9]
T. L. Hayes, G. P. Krishnan, M. Bazhenov, H. T. Siegelmann, T. J. Sejnowski, C. Kanan, Replay in deep learning: Current approaches and missing biological elements, Neural computation 33 (11) (2021) 2908–2950. 2
2021
-
[10]
Bonicelli, M
L. Bonicelli, M. Boschini, A. Porrello, C. Spampinato, S. Calderara, On the effectiveness of lipschitz-driven rehearsal in continual learning, Adv. Neural Inform. Process. Syst. 35 (2022) 31886–31901. 2
2022
-
[11]
Rebuffi, A
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, C. H. Lampert, iCaRL: Incremental Classifier and Representation Learn- ing, in: IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 2001–2010. 2, 3, 16, 19, 20, 21, 22, 23, 24, 25, 26, 28
2017
-
[12]
H. Shin, J. K. Lee, J. Kim, J. Kim, Continual Learning with Deep Generative Replay, in: Adv. Neural Inform. Process. Syst., 2017, pp. 1–10. 2
2017
-
[13]
G. M. Van de Ven, H. T. Siegelmann, A. S. Tolias, Brain-inspired replay for continual learning with artificial neural networks, Nature Communications 11 (1) (2020) 4069. 2, 16
2020
-
[14]
Mundt, I
M. Mundt, I. Pliushch, S. Majumder, Y . Hong, V . Ramesh, Unified probabilistic deep continual learning through generative replay and open set recognition, Journal of Imaging 8 (4) (2022) 93. 2
2022
-
[15]
D. Kim, B. Han, On the stability-plasticity dilemma of class-incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 20196–20204. 2, 16
2023
-
[16]
G. Yang, F. Pan, W.-B. Gan, Stably maintained dendritic spines are associated with lifelong memories, Nature 462 (7275) (2009) 920–924. 2
2009
-
[17]
D. Ji, M. A. Wilson, Coordinated memory replay in the visual cortex and hippocampus during sleep, Nature neuroscience 10 (1) (2007) 100–107. 2
2007
-
[18]
O. C. Gonzalez, Y . Sokolov, G. P. Krishnan, J. E. Delanois, M. Bazhenov, Can sleep protect memories from catastrophic forgetting?, eLife 9 (2020) e51005. 2
2020
-
[19]
D. J. Bridge, J. L. V oss, Hippocampal binding of novel information with dominant memory traces can support both memory stability and change, Journal of Neuroscience 34 (6) (2014) 2203–2213. 2 36
2014
-
[20]
Mahmoodi, M
L. Mahmoodi, M. Harandi, P. Moghadam, Flashback for continual learning, in: Proceedings of Int. Conf. Com- put. Vis., 2023, pp. 3434–3443. 3
2023
-
[21]
S. Hou, X. Pan, C. C. Loy, Z. Wang, D. Lin, Learning a Unified Classifier Incrementally via Rebalancing, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 831–839. 3, 5, 10, 16, 19, 20, 21, 22, 24, 25, 26, 28, 30
2019
-
[22]
Boschini, L
M. Boschini, L. Bonicelli, P. Buzzega, A. Porrello, S. Calderara, Class-incremental continual learning into the extended der-verse, IEEE Trans. Pattern Anal. Mach. Intell. 45 (5) (2022) 5497–5512. 3, 16, 19, 21, 22, 24
2022
-
[23]
Wang, D.-W
F.-Y . Wang, D.-W. Zhou, H.-J. Ye, D.-C. Zhan, Foster: Feature boosting and compression for class-incremental learning, in: Eur. Conf. Comput. Vis., Springer, 2022, pp. 398–414. 3, 7, 10, 17, 19, 20, 22, 23
2022
-
[24]
Wang, D.-W
F.-Y . Wang, D.-W. Zhou, L. Liu, H.-J. Ye, Y . Bian, D.-C. Zhan, P. Zhao, Beef: Bi-compatible class-incremental learning via energy-based expansion and fusion, in: Int. Conf. Learn. Represent., 2023, pp. 1–25. 3, 18, 19, 20, 22, 23
2023
-
[25]
Douillard, M
A. Douillard, M. Cord, C. Ollion, T. Robert, E. Valle, PODNet: Pooled outputs distillation for small-tasks incremental learning, in: Eur. Conf. Comput. Vis., 2020, pp. 86–102. 5, 16, 19, 23
2020
-
[26]
Buzzega, M
P. Buzzega, M. Boschini, A. Porrello, D. Abati, S. Calderara, Dark experience for general continual learning: a strong, simple baseline, Adv. Neural Inform. Process. Syst. 33 (2020) 15920–15930. 6, 10, 16, 19, 21, 22, 26
2020
-
[27]
Iscen, J
A. Iscen, J. Zhang, S. Lazebnik, C. Schmid, Memory-efficient incremental learning through feature adaptation, in: Eur. Conf. Comput. Vis., 2020, pp. 699–715. 6, 16
2020
-
[28]
Kirkpatrick, R
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ra- malho, A. Grabska-Barwinska, et al., Overcoming catastrophic forgetting in neural networks, Proceedings of the national academy of sciences 114 (13) (2017) 3521–3526. 6, ...
2017
-
[29]
S. Yan, J. Xie, X. He, DER: Dynamically Expandable Representation for Class Incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 3014–3023. 7, 17, 23
2021
-
[30]
Mundt, Y
M. Mundt, Y . Hong, I. Pliushch, V . Ramesh, A wholistic view of continual learning with deep neural networks: Forgotten lessons and the bridge to active and open world learning, Neural Networks 160 (2023) 306–336. 15
2023
-
[31]
K. Roy, C. Simon, P. Moghadam, M. Harandi, Cl3: Generalization of contrastive loss for lifelong learning, Journal of Imaging 9 (12) (2023). 16
2023
-
[32]
K. Roy, P. Moghadam, M. Harandi, L3dmc: Lifelong learning using distillation via mixed-curvature space, in: Int. Conf. on Med. Image Computing and Computer-Assisted Intervention, 2023, pp. 123–133. 16
2023
-
[33]
Simon, P
C. Simon, P. Koniusz, M. Harandi, On Learning the Geodesic Path for Incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 1591–1600. 16, 19
2021
-
[34]
M. Kang, J. Park, B. Han, Class-incremental learning by knowledge distillation with adaptive feature consolida- tion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16071–16080. 16, 19
2022
-
[35]
Zhu, X.-Y
F. Zhu, X.-Y . Zhang, C. Wang, F. Yin, C.-L. Liu, Prototype augmentation and self-supervision for incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 5871–5880. 16
2021
-
[36]
Knights, P
J. Knights, P. Moghadam, M. Ramezani, S. Sridharan, C. Fookes, Incloud: Incremental learning for point cloud place recognition, in: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2022, pp. 8559–8566. 16
2022
-
[37]
H. Yin, P. Molchanov, J. M. Alvarez, Z. Li, A. Mallya, D. Hoiem, N. K. Jha, J. Kautz, Dreaming to distill: Data- free knowledge transfer via deepinversion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 8715–8724. 16 37
2020
-
[38]
X. Li, S. Wang, J. Sun, Z. Xu, Memory efficient data-free distillation for continual learning, Pattern Recognition 144 (2023) 109875. 16
2023
-
[39]
E. Fini, S. Lathuiliere, E. Sangineto, M. Nabi, E. Ricci, Online continual learning under extreme memory con- straints, in: Eur. Conf. Comput. Vis., Springer, 2020, pp. 720–735. 16
2020
-
[40]
E. Fini, V . G. T. Da Costa, X. Alameda-Pineda, E. Ricci, K. Alahari, J. Mairal, Self-supervised models are continual learners, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 9621–9630. 16
2022
-
[41]
Gomez-Villa, B
A. Gomez-Villa, B. Twardowski, L. Yu, A. D. Bagdanov, J. van de Weijer, Continually learning self-supervised representations with projected functional regularization, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3867–3877. 16
2022
-
[42]
Q. Gu, D. Shim, F. Shkurti, Preserving linear separability in continual learning by backward feature projection, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 24286–24295. 16, 19
2023
-
[43]
Riemer, I
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y . Tu, , G. Tesauro, Learning to learn without forgetting by maximizing transfer and minimizing interference, in: Int. Conf. Learn. Represent., 2019, pp. 1–31. 16, 23
2019
-
[44]
Caccia, R
L. Caccia, R. Aljundi, N. Asadi, T. Tuytelaars, J. Pineau, E. Belilovsky, New insights on reducing abrupt repre- sentation change in online continual learning, in: Int. Conf. Learn. Represent., 2022, pp. 1–27. 16
2022
-
[45]
J. Bang, H. Kim, Y . Yoo, J.-W. Ha, J. Choi, Rainbow memory: Continual learning with a memory of diverse samples, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 8218–8227. 16
2021
-
[46]
Y . Liu, Y . Su, A.-A. Liu, B. Schiele, Q. Sun, Mnemonics Training: Multi-Class Incremental Learning without Forgetting, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 12245–12254. 16
2020
-
[47]
S. Ho, M. Liu, L. Du, L. Gao, Y . Xiang, Prototype-guided memory replay for continual learning, IEEE Trans. Neural Net. Learn. Sys. (2023). 16
2023
-
[48]
Aljundi, K
R. Aljundi, K. Kelchtermans, T. Tuytelaars, Task-free continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 11254–11263. 16
2019
-
[49]
Aljundi, E
R. Aljundi, E. Belilovsky, T. Tuytelaars, L. Charlin, M. Caccia, M. Lin, L. Page-Caccia, Online continual learn- ing with maximal interfered retrieval, Adv. Neural Inform. Process. Syst. 32 (2019). 16
2019
-
[50]
Caccia, E
L. Caccia, E. Belilovsky, M. Caccia, J. Pineau, Online learned continual compression with adaptive quantization modules, in: Int. Conf. Mach. Learn., PMLR, 2020, pp. 1240–1250. 16
2020
-
[51]
Y . Oh, D. Baek, B. Ham, Alife: Adaptive logit regularizer and feature replay for incremental semantic segmen- tation, Adv. Neural Inform. Process. Syst. 35 (2022) 14516–14528. 16
2022
-
[52]
L. Wang, X. Zhang, K. Yang, L. Yu, C. Li, L. HONG, S. Zhang, Z. Li, Y . Zhong, J. Zhu, Memory replay with data compression for continual learning, in: Int. Conf. Learn. Represent., 2022, pp. 1–25. 16
2022
-
[53]
X. Liu, C. Wu, M. Menta, L. Herranz, B. Raducanu, A. D. Bagdanov, S. Jui, J. v. de Weijer, Generative feature replay for class-incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 226–227. 16
2020
-
[54]
F. Ye, A. G. Bors, Learning latent representations across multiple data domains using lifelong vaegan, in: Eur. Conf. Comput. Vis., 2020, pp. 777–795. 16
2020
-
[55]
Binici, S
K. Binici, S. Aggarwal, N. T. Pham, K. Leman, T. Mitra, Robust and resource-efficient data-free knowledge distillation by generative pseudo replay, in: AAAI, V ol. 36, 2022, pp. 6089–6096. 16
2022
-
[56]
G. M. Van De Ven, Z. Li, A. S. Tolias, Class-incremental learning with generative classifiers, in: IEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 3611–3620. 16 38
2021
-
[57]
Y . Cong, M. Zhao, J. Li, S. Wang, L. Carin, Gan memory with no forgetting, Adv. Neural Inform. Process. Syst. 33 (2020) 16481–16494. 16
2020
-
[58]
T. Li, G. Pang, X. Bai, J. Zheng, L. Zhou, X. Ning, Learning adversarial semantic embeddings for zero-shot recognition in open worlds, Pattern Recognition 149 (2024) 110258. 16
2024
-
[59]
Lopez-Paz, M
D. Lopez-Paz, M. Ranzato, Gradient Episodic Memory for Continual Learning, in: Adv. Neural Inform. Process. Syst., 2017, pp. 1–10. 16, 20
2017
-
[60]
Chaudhry, M
A. Chaudhry, M. Ranzato, M. Rohrbach, M. Elhoseiny, Efficient lifelong learning with a-GEM, in: Int. Conf. Learn. Represent., 2019, pp. 1–20. 16
2019
-
[61]
S. Jung, H. Ahn, S. Cha, T. Moon, Continual learning with node-importance based adaptive group sparse regu- larization, Adv. Neural Inform. Process. Syst. 33 (2020) 3647–3658. 17
2020
-
[62]
J. Lee, H. G. Hong, D. Joo, J. Kim, Continual learning with extended kronecker-factored approximate curvature, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9001–9010. 17
2020
-
[63]
Zenke, B
F. Zenke, B. Poole, S. Ganguli, Continual Learning Through Synaptic Intelligence, in: Int. Conf. Mach. Learn., 2017, pp. 1–9. 17, 19
2017
-
[64]
Aljundi, F
R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, T. Tuytelaars, Memory Aware Synapses: Learning what (not) to forget, in: Eur. Conf. Comput. Vis., 2018, pp. 1–16. 17
2018
-
[65]
H. Ran, W. Li, L. Li, S. Tian, X. Ning, P. Tiwari, Learning optimal inter-class margin adaptively for few- shot class-incremental learning via neural collapse-based meta-learning, Information Processing & Management 61 (3) (2024) 103664. 17
2024
-
[66]
Pelosin, S
F. Pelosin, S. Jha, A. Torsello, B. Raducanu, J. van de Weijer, Towards exemplar-free continual learning in vision transformers: an account of attention, functional and weight regularization, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 3820–3829. 17
2022
-
[67]
Soutif-Cormerais, A
A. Soutif-Cormerais, A. Carta, J. van de Weijer, Improving online continual learning performance and stability with temporal ensembles, in: S. Chandar, R. Pascanu, H. Sedghi, D. Precup (Eds.), Proceedings of Machine Learn. Research, V ol. 232, PMLR, 2023, pp. 828–845. 17
2023
-
[68]
Gupta, K
G. Gupta, K. Yadav, L. Paull, La-maml: Look-ahead meta learning for continual learning. 2020, URL https://arxiv. org/abs (2007). 17
2007
-
[69]
L. Wang, M. Zhang, Z. Jia, Q. Li, C. Bao, K. Ma, J. Zhu, Y . Zhong, Afec: Active forgetting of negative transfer in continual learning, Adv. Neural Inform. Process. Syst. 34 (2021) 22379–22391. 18
2021
-
[70]
L. Wang, X. Zhang, Q. Li, M. Zhang, H. Su, J. Zhu, Y . Zhong, Incorporating neuro-inspired adaptability for continual learning in artificial intelligence, Nature Machine Intelligence 5 (12) (2023) 1356–1368. 18
2023
-
[71]
Serra, D
J. Serra, D. Suris, M. Miron, A. Karatzoglou, Overcoming catastrophic forgetting with hard attention to the task, in: Int. Conf. Learn. Represent., PMLR, 2018, pp. 4548–4557. 17
2018
-
[72]
Ostapenko, M
O. Ostapenko, M. Puscas, T. Klein, P. Jahnichen, M. Nabi, Learning to Remember: A Synaptic Plasticity Driven Framework for Continual Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 11313–11321. 17
2019
-
[73]
J. Yoon, E. Yang, J. Lee, S. J. Hwang, Lifelong Learning with Dynamically Expandable Networks, in: Int. Conf. Learn. Represent., 2018, pp. 1–11. 17
2018
-
[74]
X. Li, Y . Zhou, T. Wu, R. Socher, C. Xiong, Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting, in: Int. Conf. Mach. Learn., PMLR, 2019, pp. 3925–3934. 17 39
2019
-
[75]
G. Lin, H. Chu, H. Lai, Towards better plasticity-stability trade-off in incremental learning: A simple linear connector, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 89–98. 18
2022
-
[76]
Krizhevsky, G
A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, https://www.cs.toronto.edu/ kriz/learning-features-2009-TR.pdf (2009) 3–58. 19
2009
-
[77]
J. Wu, Q. Zhang, G. Xu, Tiny imagenet challenge, Technical Report (2017). 19
2017
-
[78]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: IEEE Conf. Comput. Vis. Pattern Recog., Ieee, 2009, pp. 248–255. 19
2009
-
[79]
Y . Wu, Y . Chen, L. Wang, Y . Ye, Z. Liu, Y . Guo, Y . Fu, Large scale incremental learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 374–382. 19, 22, 23, 26
2019
-
[80]
S. Kim, L. Noci, A. Orvieto, T. Hofmann, Achieving a better stability-plasticity trade-off via auxiliary networks in continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 11930–11939. 20, 24
2023
-
[81]
Y . Liu, B. Schiele, Q. Sun, Adaptive aggregation networks for class-incremental learning, in: IEEE Conf. Com- put. Vis. Pattern Recog., 2021, pp. 2544–2553. 20, 24
2021
-
[82]
K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 770–778. 22, 23
2016
-
[83]
F. M. Castro, M. J. Marín-Jiménez, N. Guil, C. Schmid, K. Alahari, End-to-End Incremental Learning, in: Eur. Conf. Comput. Vis., 2018, pp. 1–16. 22, 26
2018
-
[84]
H. Ahn, J. Kwak, S. Lim, H. Bang, H. Kim, T. Moon, SS-IL: Separated Softmax for Incremental Learning, in: Int. Conf. Comput. Vis., 2021, pp. 844–853. 22
2021
-
[85]
B. Zhao, X. Xiao, G. Gan, B. Zhang, S.-T. Xia, Maintaining Discrimination and Fairness in Class incremental Learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 13208–13207. 23
2020
-
[86]
Douillard, A
A. Douillard, A. Ramé, G. Couairon, M. Cord, Dytox: Transformers for continual learning with dynamic token expansion, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 9285–9295. 23
2022
-
[87]
Dosovitskiy, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)
A. Dosovitskiy, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020). 26
2020 arXiv
-
[88]
Kornblith, M
S. Kornblith, M. Norouzi, H. Lee, G. Hinton, Similarity of neural network representations revisited, in: Int. Conf. Mach. Learn., PMLR, 2019, pp. 3519–3529. 27
2019
-
[89]
Yosinski, J
J. Yosinski, J. Clune, Y . Bengio, H. Lipson, How transferable are features in deep neural networks?, Adv. Neural Inform. Process. Syst. 27 (2014). 28
2014
-
[90]
Z. Wang, Z. Zhang, C.-Y . Lee, H. Zhang, R. Sun, X. Ren, G. Su, V . Perot, J. Dy, T. Pfister, Learning to prompt for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 139–149. 28
2022
-
[91]
J. S. Smith, L. Karlinsky, V . Gutta, P. Cascante-Bonilla, D. Kim, A. Arbelle, R. Panda, R. Feris, Z. Kira, Coda- prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 11909–11919. 28
2023
-
[92]
Liang, W.-J
Y .-S. Liang, W.-J. Li, Inflora: Interference-free low-rank adaptation for continual learning, in: IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 23638–23647. 28
2024
-
[93]
X. Wang, T. Chen, Q. Ge, H. Xia, R. Bao, R. Zheng, Q. Zhang, T. Gui, X.-J. Huang, Orthogonal subspace learning for language model continual learning, in: Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 10658–10671. 28
2023
-
[94]
C. M. Bishop, H. Bishop, Deep learning: Foundations and concepts, Springer Nature, 2023. 33 40
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.