Pith. sign in

REVIEW 5 major objections 5 minor 198 references

Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims to be the first systematic literature review of Online Continual Learning, mapping 81 approaches, more than 500 components, over 1,000 features, and 83 datasets into a structured synthesis.

desk verdict A genuinely useful compilation of OCL approaches and datasets, but the unvalidated screening filter and internal count inconsistencies weaken the comprehensiveness claim. read the letter →

arxiv 2501.04897 v1 pith:IONT3QVP submitted 2025-01-09 cs.LG

classification cs.LG
keywords OnlineContinualLearningCatastrophicForgettingIncrementalLifelongStability-PlasticityTrade-offSystematicLiteratureReviewReplay-basedMethodsBenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish the first systematic literature review of Online Continual Learning (OCL), the subfield of machine learning where models must keep learning from data streams in real time without forgetting earlier knowledge. It argues that prior reviews were too narrow, and that a guideline-driven systematic review can produce a comprehensive map of 81 approaches, more than 500 components, over 1,000 features, and 83 datasets. If correct, researchers gain a common vocabulary, a structured taxonomy of replay-, architecture-, and regularization-based methods, and an evidence-based list of open problems such as computational overhead, domain-agnostic solutions, and scalability. The review also proposes a unified definition of OCL and identifies underexplored directions like self-supervised multimodal learning and adaptive memory combining sparse retrieval with generative replay.

What carries the argument

The central machinery is the systematic literature review protocol the authors adopt, executed through four phases: building a pool of 2,061 publications, applying inclusion and exclusion criteria, assessing quality, and extracting data. Relevance is decided by semantic similarity between each paper and two keyword sets using a sentence-embedding model, with a fixed 0.5 threshold separating papers into high, medium, and low relevance; high-relevance papers that answer at least three of five quality questions affirmatively proceed to extraction. This protocol converts a diffuse literature into countable entities—approaches, components, features, datasets, and quality attributes—and the frequency of these entities is what supports the review's conclusions.

What would settle it

Run the same relevance screening on the pooled 2,061 papers with the similarity threshold moved from 0.5 to, say, 0.3 and 0.7, and compare the resulting high-relevance sets: if a major, widely cited OCL approach falls out at the stricter threshold or enters only at the looser one, the completeness of the 81-paper corpus and the counts built from it are not stable.

Watch

Extended reading notes

Core claim

The central claim is that Online Continual Learning had no comprehensive, methodologically rigorous synthesis, and that this review supplies one. The authors analyze 81 unique OCL approaches, categorize them into three main strategy types, extract more than 500 components and over 1,000 features, and compile a list of 83 datasets. They use frequency counts of these entities to identify dominant trends, such as the prevalence of replay-based methods, and to surface under-researched areas such as non-visual and multimodal tasks. The review also offers a unified definition of OCL based on three conditions: small real-time batches, disjoint label spaces across tasks, and single-epoch training with no revisiting of past data.

Load-bearing premise

The load-bearing premise is that the automatic relevance screening, which uses sentence-embedding similarity with a fixed 0.5 threshold to label papers as high, medium, or low relevance, correctly captures the OCL literature so that the final 81 papers are representative of the whole field.

Editorial extensions

If this is right

  • Researchers gain a shared reference for OCL definitions and settings, which can reduce the terminological duplication noted in the field.
  • Because replay-based methods appear in 62 of 81 approaches, future method design and benchmarking should treat memory-buffer and generative-replay baselines as the default comparison points.
  • The 83-dataset inventory shows a heavy concentration on image classification and reveals that audio, text, time-series, and multimodal continual learning remain comparatively underexplored.
  • The review's list of unresolved challenges, including computational overhead, domain-agnostic solutions, and scalability, provides a concrete research agenda for the next generation of OCL methods.
  • Hybrid approaches that combine sparse retrieval with generative replay, and self-supervised learning for multimodal data, are identified as promising directions worth prioritizing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable next step is to validate the semantic-similarity screening by manually labeling a random sample of the 2,061-paper pool and measuring agreement with the 0.5-threshold labels.
  • Because the review counts prevalence rather than empirical success, its conclusion that replay-based methods dominate should be read as a statement about research activity, not about which strategy performs best; a performance-oriented meta-analysis would require standardized benchmarks.
  • The unified OCL definition could seed a shared evaluation protocol: fixing the three conditions of small real-time batches, disjoint label spaces, and single-epoch training would make results across future papers directly comparable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper presents a systematic literature review (SLR) of online continual learning (OCL), following Newman's guidelines. The authors describe a multi-phase process: pooling 2,061 publications, applying keyword-based and semantic-similarity filtering, quality assessment, and data extraction, leading to an analysis of 81 OCL approaches. The review categorizes approaches into replay-, architecture-, and regularization-based strategies; reports components, features, quality attributes, and datasets; and discusses open challenges and future directions. The paper claims to be the first SLR in OCL and provides a GitHub link for the complete methodology and extracted data.

Significance. If the reported counts and corpus were reliable, the paper would be a useful reference, assembling a structured overview of OCL approaches, datasets, components, and metrics, with a documented methodology and public materials. Strengths include the explicit adoption of a named SLR framework, a pipeline with publication counts at each phase, and the effort to map approaches to components, features, and datasets in appendices. However, the reliability of the synthesis is currently compromised by an unvalidated semantic-similarity threshold at the selection stage and by major internal inconsistencies in the headline numbers. These issues affect the central claim of comprehensiveness, so the paper needs substantial revision before it can serve as the reference it aims to be.

major comments (5)
  1. [2.3] The semantic similarity screening uses Sentence-BERT with a threshold of 0.5 to categorize papers as high, medium, or low relevance, but no calibration, precision/recall evaluation, inter-annotator agreement, or manual audit of excluded papers is reported. Because this step determines the 170 high-relevance papers and ultimately the 81 approaches analyzed, all downstream counts inherit any bias introduced here; the authors should provide a recall audit against a gold-standard set of OCL papers and a sensitivity analysis of the chosen threshold.
  2. [2.4] The refined search string is built from the same Google Scholar 'Initial Hypothesis' set that also yields the keyword sets K1 and K2; this circularity means any bias in the top-227 Google Scholar results propagates into both the semantic filter and the final search strategy. The authors should validate their search string independently, for example by comparing it with a search string developed by domain experts or by measuring how many well-known OCL papers are missed.
  3. [Abstract, 4.2, and Appendices B-E] The abstract claims 'over 1,000 features' and 'more than 500 components,' while Section 4.2 reports '127 components, 48 features' and Section 3.5 states '51 key quality attributes' (Section 4.2 then says '60 quality attributes'). These are not cosmetic discrepancies: the paper's central contribution is the quantitative characterization of the field, and the tables in Appendices B-E show many approaches with zero extracted components, features, or datasets (e.g., Tables 3, 5, and 7). The authors must reconcile the counts, clearly distinguish between unique entity types and total occurrences, and complete or explicitly annotate the sparse tables before the comprehensiveness claim can be evaluated.
  4. [2.7] The quality-assessment step applies five yes/no questions and retains papers with at least three 'yes' answers, but the questions are subjective ('clear problem statement,' 'research challenge are well-defined') and no inter-rater reliability or pilot validation is reported. Since this step further filters the set from 170 to 81 papers, the authors should document the number of reviewers, the agreement rate, and a sensitivity analysis of the cutoff.
  5. [Appendix E and 4.4] The paper states that the complete list of extracted features can be found in the appendix and that the mapping tables show how components are combined, but the appendix tables are heavily incomplete; for example, Table 7 lists zero datasets for multiple approaches, and Section 4.4 admits that a component-feature mapping table was omitted because of 'significant sparsity in the mapping table.' Incomplete extraction tables prevent readers from verifying the aggregate counts and undermine the reproducibility of the review; the authors should either complete the tables or clearly report coverage rates per approach and per entity.
minor comments (5)
  1. [Abstract] The GitHub URL in the abstract contains a space ('kiyan-rezaee/ Systematic-Literature-Review...'); it should be a single clickable link and the repository should be verified to contain the promised artifacts.
  2. [Figure 7] Figure 7 appears to be a screenshot of a presentation slide with interface text such as 'Share Made with...'; the authors should replace it with a clean, publication-quality figure.
  3. [3.3.1] The claim that ResNet18 appears in '62 out of 81 approaches' is not supported by the check marks in Table 3; please reconcile the frequency count with the appendix table.
  4. [2.6 and 3.3.1] There are typographical spacing issues in the text, such as 'F eatures' in Section 2.6 and 'V ariational Autoencoders' in Section 3.3.1; a careful proofread of the manuscript is needed.
  5. [4.3] Section 4.3 acknowledges that the frequency-based approach 'lacks the depth required to delve into the theoretical underpinnings,' but this limitation should be mentioned earlier in the introduction or methodology so that readers are not misled by the comprehensiveness claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found; the review synthesizes external literature, and the unvalidated relevance filter is a validity risk, not a circular derivation.

full rationale

This paper is a systematic literature review, so its outputs (taxonomies, counts, and trends) are summaries of the external papers it reviews rather than derivations from fitted parameters or self-referential definitions. The methodology chain described in Section 2—keyword extraction, Sentence-BERT relevance filtering, quality assessment, and entity coding—does not define any output in terms of the paper's own conclusions, and no fitted value is later renamed as a prediction. The Sentence-BERT threshold of 0.5 in Section 2.3 is uncalibrated, which threatens the representativeness and therefore the strength of the 'comprehensive' claim, but that is a selection-bias and validity concern, not circularity: the retained paper set is not constructed to force the paper's findings. The review also does not rely on load-bearing self-citations; its methodological authorities (Newman's guidelines [45], Sentence-BERT [172]) are external. The numerical inconsistencies between the abstract ('over 1,000 features', 'more than 500 components') and Section 4.2 ('127 components, 48 features, 60 quality attributes'), as well as the Appendix C total of 558 feature occurrences, are correctness and reporting issues rather than reductions of outputs to inputs. Accordingly, no circular step can be exhibited, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a review paper; there are no novel entities or fitted parameters. The central claim rests on methodological assumptions about literature screening, quality assessment, and coding consistency.

assumptions (3)
  • domain assumption The search strategy and semantic similarity threshold of 0.5 correctly identify all high-relevance OCL papers.
    The review's completeness depends on the screening step, but the threshold is not validated against a gold standard, and the keyword sets K1/K2 are not fully specified.
  • domain assumption The five quality assessment questions with a 'three yes' threshold sufficiently distinguish high-quality from low-quality studies.
    This binary heuristic is arbitrary and may include or exclude papers inconsistently.
  • domain assumption Manual coding of entities (approaches, components, features, datasets) was performed consistently across all 81 papers.
    No inter-coder reliability is reported, and the sparse appendix tables suggest possible extraction errors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks." pith.science (2026). https://pith.science/paper/IONT3QVP

@misc{pith2026250104897,
  author       = {Pith},
  title        = {Pith review of: Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IONT3QVP}},
  note         = {Machine review of arXiv:2501.04897}
}
read the original abstract

Online Continual Learning (OCL) is a critical area in machine learning, focusing on enabling models to adapt to evolving data streams in real-time while addressing challenges such as catastrophic forgetting and the stability-plasticity trade-off. This study conducts the first comprehensive Systematic Literature Review (SLR) on OCL, analyzing 81 approaches, extracting over 1,000 features (specific tasks addressed by these approaches), and identifying more than 500 components (sub-models within approaches, including algorithms and tools). We also review 83 datasets spanning applications like image classification, object detection, and multimodal vision-language tasks. Our findings highlight key challenges, including reducing computational overhead, developing domain-agnostic solutions, and improving scalability in resource-constrained environments. Furthermore, we identify promising directions for future research, such as leveraging self-supervised learning for multimodal and sequential data, designing adaptive memory mechanisms that integrate sparse retrieval and generative replay, and creating efficient frameworks for real-world applications with noisy or evolving task boundaries. By providing a rigorous and structured synthesis of the current state of OCL, this review offers a valuable resource for advancing this field and addressing its critical challenges and opportunities. The complete SLR methodology steps and extracted data are publicly available through the provided link: https://github.com/kiyan-rezaee/ Systematic-Literature-Review-on-Online-Continual-Learning

Figures

Figures reproduced from arXiv: 2501.04897 by the authors.

Figure 1
Figure 1. Our systematic methodology in this work. Our methodology starts with (1) defining research objectives, followed by developing research questions. (2) A conceptual framework is designed, and (3) selection criteria such as relevancy are constructed. (4) A search strategy is implemented using initial keywords and refined to extract more publications. (5) Studies are selected by applying the criteria to identify highly … view at source ↗
Figure 2
Figure 2. Distribution of articles extracted from multiple digital libraries. This chart depicts the number of [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Classification of OCL approaches by strategy. The figure categorizes 81 identified approaches into [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Overview of key strategies in OCL. This figure illustrates three primary strategies in OCL: (1) [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Prevalent components in OCL literature. The figure highlights frequently used components across [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Key features in OCL literature. This figure highlights the most important and prevalent features [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Prominent datasets in OCL literature. This figure categorizes and lists key datasets employed in [PITH_FULL_IMAGE:figures/full_fig_p022_7.png]
Figure 8
Figure 8. Figure 8: Variants of datasets in OCL literature. This figure visualizes the relationships between major [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

198 extracted references · 29 canonical work pages

  1. [1]

    A Comprehensive Survey of Continual Learning: Theory, Method and Application , L. Wang, X. Zhang, H. Su, and J. Zhu. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5362-5383, 2024. [Online]. Available: https://doi.org/10.1109/TPAMI.2024.3367329

  2. [2]

    Lesort, V

    Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges , T. Lesort, V. Lomonaco, A. Stoian, D. Maltoni, D. Filliat, and N. D ´ ıaz-Rodr ´ ıguez. Information Fusion, vol. 58, pp. 52–68, 2020. [Online]. Available: https://doi.org/10.1016/j.inffus.2019.12.004

  3. [3]

    Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare , J. Feng, R. V. Phillips, I. Malenica, A. Bishara, A. E. Hubbard, L. A. Celi, and R. Pirracchio. NPJ Digital Medicine, vol. 5, no. 1, p. 66, 2022. [Online]. Available: https://doi.org/10.1038/s41746-022-00622-y

  4. [4]

    Online Continual Learning For Interactive Instruction Following Agents , B. Kim, M. Seo, and J. Choi. arXiv preprint arXiv:2403.07548, 2024. [Online]. Available: https://arxiv.org/abs/2403.07548

  5. [5]

    Learning Equi-angular Representations for Online Continual Learning , M. Seo, H. Koh, W. Jeung, M. Lee, S. Kim, H. Lee, S. Cho, S. Choi, H. Kim, and J. Choi. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23933–23942, 2024

  6. [6]

    Improving Plasticity in Online Continual Learning via Collaborative Learning , M. Wang, N. Michel, L. Xiao, and T. Yamasaki. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23460–23469, 2024

  7. [7]

    Raghavan, J

    DELTA: Decoupling Long-Tailed Online Continual Learning, S. Raghavan, J. He, and F. Zhu. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4054–4064, 2024

  8. [8]

    Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-Distillation , H. Yan, L. Wang, K. Ma, and Y. Zhong. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23670–23680, 2024

Show all 198 references
  1. [9]

    Online-LoRA: Task-free Online Continual Learning via Low Rank Adaptation , X. Wei, G. Li, and R. Marculescu. arXiv preprint arXiv:2411.05663, 2024. [Online]. Available: https://arxiv.org/abs/2411.05663

  2. [10]

    Progressive Prototype Evolving for Dual-Forgetting Mitigation in Non-Exemplar Online Continual Learning, Q. Li, Y. Peng, and J. Zhou. Proceedings of the 32nd ACM International Conference on Multimedia, pp. 2477–2486, 2024. 29

  3. [11]

    Adaptive Shortcut Debiasing for Online Continual Learning , D. Kim, D. Park, Y. Shin, J. Bang, H. Song, and J. Lee. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 12, pp. 13122–13131, 2024

  4. [12]

    Overcoming Domain Drift in Online Continual Learning , F. Lyu, D. Liu, L. Zhao, Z. Zhang, F. Shang, F. Hu, W. Feng, and L. Wang. arXiv preprint arXiv:2405.09133, 2024. [Online]. Available: https://arxiv.org/abs/2405.09133

  5. [13]

    Vander Eeckt and others

    Unsupervised Online Continual Learning for Automatic Speech Recognition , S. Vander Eeckt and others. arXiv preprint arXiv:2406.12503, 2024. [Online]. Available: https://arxiv.org/abs/2406.12503

  6. [14]

    Forgetting, Ignorance or Myopia: Revisiting Key Challenges in Online Continual Learning , X. Wang, C. Geng, W. Wan, S.-y. Li, and S. Chen. arXiv preprint arXiv:2409.19245, 2024. [Online]. Available: https://arxiv.org/abs/2409.19245

  7. [15]

    SRTFD: Scalable Real-Time Fault Diagnosis through Online Continual Learning , D. Zhao, K. Sharma, H. Yin, Y. Qi, and S. Zhang. arXiv preprint arXiv:2408.05681, 2024. [Online]. Available: https://arxiv.org/abs/2408.05681

  8. [16]

    ER-FSL: Experience Replay with Feature Subspace Learning for Online Continual Learning , H. Lin. arXiv preprint arXiv:2407.12279, 2024. [Online]. Available: https://arxiv.org/abs/2407.12279

  9. [17]

    Dual-CBA: Improving Online Continual Learning via Dual Continual Bias Adaptors from a Bi-level Optimization Perspective, Q. Wang, R. Wang, Y. Wu, X. Jia, M. Zhou, and D. Meng. arXiv preprint arXiv:2408.13991, 2024. [Online]. Available: https://arxiv.org/abs/2408.13991

  10. [18]

    Cheng, Y

    NLOCL: Noise-Labeled Online Continual Learning , K. Cheng, Y. Ma, G. Wang, L. Zong, and X. Liu. Electronics, vol. 13, no. 13, p. 2560, 2024

  11. [19]

    Adaptive VIO: Deep Visual-Inertial Odometry with Online Continual Learning , Y. Pan, W. Zhou, Y. Cao, and H. Zha. arXiv preprint arXiv:2405.16754, 2024. [Online]. Available: https://arxiv.org/abs/2405.16754

  12. [20]

    Caccia, E

    Online learned continual compression with adaptive quantization modules , L. Caccia, E. Belilovsky, M. Caccia, and J. Pineau. International Conference on Machine Learning, pp. 1240–1250, 2020. [Online]. Available: https://proceedings.mlr.press

  13. [21]

    Chrysakis and M.-F

    Online continual learning from imbalanced data, A. Chrysakis and M.-F. Moens. International Conference on Machine Learning, pp. 1952–1961, 2020. [Online]. Available: https://proceedings.mlr.press

  14. [22]

    Gupta, K

    Look-ahead meta learning for continual learning , G. Gupta, K. Yadav, and L. Paull. Advances in Neural Information Processing Systems, vol. 33, pp. 11588–11598, 2020. [Online]. Available: https://proceedings.neurips.cc

  15. [24]

    Gradient-based editing of memory examples for online task-free continual learning , X. Jin, A. Sadhu, J. Du, and X. Ren. Advances in Neural Information Processing Systems, vol. 34, pp. 29193–29205, 2021. [Online]. Available: https://proceedings.neurips.cc

  16. [25]

    Mitigating forgetting in online continual learning with neuron calibration , H. Yin, P. Li, et al. Ad- vances in Neural Information Processing Systems, vol. 34, pp. 10260–10272, 2021. [Online]. Available: https://proceedings.neurips.cc

  17. [26]

    Continual learning on noisy data streams via self-purified replay , C. D. Kim, J. Jeong, S. Moon, and G. Kim. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 537–547, 2021. [Online]. Available: https://openaccess.thecvf.com 30

  18. [27]

    Liang and W.-J

    Optimizing class distribution in memory for multi-label online continual learning , Y.-S. Liang and W.-J. Li. arXiv preprint arXiv:2209.11469, 2022. [Online]. Available: https://arxiv.org/abs/2209.11469

  19. [28]

    Schedule-robust online continual learning , R. Wang, M. Ciccone, G. Luise, A. Yapp, M. Pontil, and C. Ciliberto. arXiv preprint arXiv:2210.05561, 2022. [Online]. Available: https://arxiv.org/abs/2210.05561

  20. [29]

    Michel, G

    Learning representations on the unit sphere: Application to online continual learning , N. Michel, G. Chierchia, R. Negrel, and J.-F. Bercher, 2023

  21. [30]

    V¨ odisch, K

    CoDEPS: Online continual learning for depth estimation and panoptic segmentation , N. V¨ odisch, K. Petek, W. Burgard, and A. Valada. arXiv preprint arXiv:2303.10147, 2023. [Online]. Available: https://arxiv.org/abs/2303.10147

  22. [31]

    Non-exemplar Online Class-Incremental Continual Learning via Dual-Prototype Self-Augment and Refinement, F. Huo, W. Xu, J. Guo, H. Wang, and Y. Fan. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 11, pp. 12698–12707, 2024. [Online]. Available: http...

  23. [32]

    Soutif- Cormerais, A

    Improving online continual learning performance and stability with temporal ensembles , A. Soutif- Cormerais, A. Carta, and J. Van de Weijer. Conference on Lifelong Learning Agents, pp. 828–845, 2023. [Online]. Available: https://proceedings.mlr.press

  24. [33]

    Summarizing stream data for memory-restricted online continual learning , J. Gu, K. Wang, W. Jiang, and Y. You. arXiv preprint arXiv:2305.16645, vol. 2, 2023. [Online]. Available: https://arxiv.org/abs/2305.16645

  25. [34]

    Wanderlust: Online continual object detection in the real world , J. Wang, X. Wang, Y. Shang-Guan, and A. Gupta. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10829–10838,

  26. [35]

    Li and D

    Learning without forgetting , Z. Li and D. Hoiem. IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, vol. 40, no. 12, pp. 2935–2947, 2017. [Online]. Available: https://doi.org/10.1109/TPAMI.2017.2773081

  27. [36]

    A Statistical Theory of Regularization-Based Continual Learning , X. Zhao, H. Wang, W. Huang, and W. Lin. arXiv preprint arXiv:2406.06213, 2024. [Online]. Available: https://arxiv.org/abs/2406.06213

  28. [37]

    Moradi, R

    A survey of regularization strategies for deep models , R. Moradi, R. Berangi, and B. Minaei. Artificial In- telligence Review, vol. 53, no. 6, pp. 3947–3986, 2020. [Online]. Available: https://doi.org/10.1007/s10462- 019-09760-3

  29. [38]

    Huang, Y

    Continual learning for text classification with information disentanglement based regularization , Y. Huang, Y. Zhang, J. Chen, X. Wang, and D. Yang. arXiv preprint arXiv:2104.05489, 2021. [Online]. Available: https://arxiv.org/abs/2104.05489

  30. [39]

    Self-improving reactive agents based on reinforcement learning, planning and teaching , L.-J. Lin. Machine Learning, vol. 8, pp. 293–321, 1992. [Online]. Available: https://doi.org/10.1007/BF00992699

  31. [40]

    Verwimp, S

    Continual learning: Applications and the road forward , E. Verwimp, S. Ben-David, M. Bethge, A. Cossu, A. Gepperth, T. L. Hayes, E. H¨ ullermeier, C. Kanan, D. Kudithipudi, C. H. Lampert, et al. arXiv preprint arXiv:2311.11908, 2023. [Online]. Available: https://arxiv.org/abs/...

  32. [41]

    Online continual learning of end-to-end speech recognition models , M. Yang, I. Lane, and S. Watanabe. arXiv preprint arXiv:2207.05071, 2022. [Online]. Available: https://arxiv.org/abs/2207.05071

  33. [42]

    Priem, H

    OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts , J. Priem, H. Piwowar, and R. Orr. arXiv preprint arXiv:2205.01833, 2022. [Online]. Available: https://arxiv.org/abs/2205.01833 31

  34. [43]

    Systematic Literature Reviews: An Introduction , G. Lame. Proceedings of the Design Society: In- ternational Conference on Engineering Design, vol. 1, no. 1, pp. 1633–1642, 2019. [Online]. Available: https://doi.org/10.1017/dsi.2019.169

  35. [44]

    Nightingale

    A guide to systematic literature reviews , A. Nightingale. Surgery (Oxford), vol. 27, pp. 381–384, 2009. [Online]. Available: https://doi.org/10.1016/j.mpsur.2009.07.005

  36. [45]

    Newman and D

    Systematic Reviews in Educational Research: Methodology, Perspectives and Application , M. Newman and D. Gough. Springer Fachmedien Wiesbaden, 2020, pp. 3–22. [Online]. Available: https://doi.org/10.1007/978-3-658-27602-7 1

  37. [46]

    Bonicelli, M

    On the effectiveness of equivariant regularization for robust online continual learning , L. Bonicelli, M. Boschini, E. Frascaroli, A. Porrello, M. Pennisi, G. Bellitto, S. Palazzo, C. Spampinato, and S. Calderara. arXiv preprint arXiv:2305.03648, 2023. [Online]. Available: ht...

  38. [47]

    Cossu, A

    Continual learning for recurrent neural networks: An empirical evaluation , A. Cossu, A. Carta, V. Lomonaco, and D. Bacciu. Neural Networks, vol. 143, pp. 607–627, 2021. [Online]. Available: https://doi.org/10.1016/j.neunet.2021.07.021

  39. [48]

    Shaheen, M

    Continual Learning for Real-World Autonomous Systems: Algorithms, Challenges and Frameworks , K. Shaheen, M. A. Hanif, O. Hasan, M. Shafique, Journal of Intelligent & Robotic Systems, vol. 105, no. 1, p. 9, Apr. 2022

  40. [49]

    Recent Advances of Foundation Language Models-based Continual Learning: A Survey , Y. Yang, J. Zhou, X. Ding, T. Huai, S. Liu, Q. Chen, L. He, Y. Xie, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2405.18653

  41. [50]

    Wickramasinghe, G

    Continual Learning: A Review of Techniques, Challenges, and Future Directions , B. Wickramasinghe, G. Saha, K. Roy, IEEE Transactions on Artificial Intelligence, vol. 5, no. 6, pp. 2526-2546, 2024. [Online]. Available: https://doi.org/10.1109/TAI.2023.3339091

  42. [51]

    Recent Advances of Continual Learning in Computer Vision: An Overview , H. Qu, H. Rahmani, L. Xu, B. Williams, J. Liu, 2021. [Online]. Available: https://doi.org/10.48550/arXiv.2109.11369

  43. [52]

    LeCun, L

    Gradient-based learning applied to document recognition , Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998. [Online]. Available: https://doi.org/10.1109/5.726791

  44. [53]

    Deep residual learning for image recognition , K. He, X. Zhang, S. Ren, J. Sun, in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. [Online]. Available: https://doi.org/10.1109/CVPR.2016.90

  45. [54]

    Revisiting Neural Networks for Continual Learning: An Architectural Perspective , A. Lu, T. Feng, H. Yuan, X. Song, Y. Sun, 2024. [Online]. Available: https://arxiv.org/abs/2404.14829

  46. [55]

    Towards Redundancy-Free Sub-networks in Continual Learning, C. Chen, J. Song, L. Gao, H. Shen, 2024. [Online]. Available: https://arxiv.org/abs/2312.00840

  47. [56]

    Continual Learning with Deep Generative Replay , H. Shin, J. K. Lee, J. Kim, J. Kim, in Advances in Neural Information Processing Systems, vol. 30, 2017

  48. [57]

    Madaan, H

    Heterogeneous Continual Learning, D. Madaan, H. Yin, W. Byeon, J. Kautz, P. Molchanov, inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 15985- 15995

  49. [58]

    Task-aware network: Mitigation of task-aware and task-free performance gap in online continual learn- ing, Y. Hong, S. Park, H. Byun, Neurocomputing, vol. 552, pp. 126527, 2023. [Online]. Available: https://doi.org/10.1016/j.neucom.2023.126527 32

  50. [59]

    Learning representations by back-propagating errors, D. E. Rumelhart, G. E. Hinton, R. J. Williams, Nature, vol. 323, no. 6088, pp. 533–536, 1986

  51. [60]

    Contextual Transformation Networks for Online Continual Learning , Q. Pham, C. Liu, D. Sa- hoo, S. HOI, in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=zx uX-BO7CH

  52. [61]

    Elharrouss, Y

    Backbones-review: Feature extractor networks for deep learning and deep reinforcement learning ap- proaches in computer vision , O. Elharrouss, Y. Akbari, N. Almadeed, S. Al-Maadeed, Computer Science Review, vol. 53, pp. 100645, 2024. [Online]. Available: http://dx.doi.org/10....

  53. [62]

    V¨ odisch, D

    CoVIO: Online Continual Learning for Visual-Inertial Odometry , N. V¨ odisch, D. Cattaneo, W. Burgard, A. Valada, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , 2023, pp. 2464-2473

  54. [63]

    Revealing the real-world applicable setting of online continual learning , Z. Xu, H. Hu, L. Liu, in 2022 IEEE 24th International Workshop on Multimedia Signal Processing (MMSP) , 2022, pp. 1-5. [Online]. Available: https://doi.org/10.1109/MMSP55362.2022.9948735

  55. [64]

    Michel, R

    Contrastive Learning for Online Semi-Supervised General Continual Learning , N. Michel, R. Negrel, G. Chierchia, J.-F. Bercher, 2022. [Online]. Available: https://arxiv.org/abs/2207.05615

  56. [65]

    Yu, W.-C

    Mitigating Forgetting in Online Continual Learning via Contrasting Semantically Distinct Augmentations , S.-F. Yu, W.-C. Chiu, 2023. [Online]. Available: https://arxiv.org/abs/2211.05347

  57. [66]

    Delange, R

    A continual learning survey: Defying forgetting in classification tasks , M. Delange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 1, pp. 1-1, 2021. [Online...

  58. [67]

    An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks , I. J. Good- fellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio. arXiv preprint arXiv:1312.6211, 2015. [Online]. Available: https://arxiv.org/abs/1312.6211

  59. [68]

    Hinton, O

    Distilling the Knowledge in a Neural Network , G. Hinton, O. Vinyals, and J. Dean. arXiv preprint arXiv:1503.02531, 2015. [Online]. Available: https://arxiv.org/abs/1503.02531

  60. [69]

    Nair and G

    Rectified linear units improve restricted boltzmann machines , V. Nair and G. E. Hinton. In Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel, 2010, pp. 807–814. ISBN: 9781605589077

  61. [70]

    Pattern Recognition and Machine Learning , C. M. Bishop. New York: Springer, 2006. ISBN: 978- 0387310732

  62. [71]

    Hendrycks and K

    Gaussian Error Linear Units (GELUs) , D. Hendrycks and K. Gimpel. In Proceedings of the 34th International Conference on Machine Learning (ICML), 2016, pp. 673–682. [Online]. Available: http://proceedings.mlr.press/v70/hendrycks17a.html

  63. [72]

    Caccia, R

    Reducing representation drift in online continual learning, L. Caccia, R. Aljundi, T. Tuytelaars, J. Pineau, and E. Belilovsky. arXiv preprint arXiv:2104.05025, vol. 1, no. 3, 2021

  64. [73]

    Online Continual Learning through Mutual Information Maximization , Y. Guo, B. Liu, and D. Zhao. In Proceedings of the 39th International Conference on Machine Learning (ICML), 2022, pp. 8109–8126. [Online]. Available: https://proceedings.mlr.press/v162/guo22g.html

  65. [74]

    Rolnick, A

    Experience Replay for Continual Learning , D. Rolnick, A. Ahuja, J. Schwarz, T. P. Lillicrap, and G. Wayne. arXiv preprint arXiv:1811.11682, 2019. [Online]. Available: https://arxiv.org/abs/1811.11682 33

  66. [75]

    Dealing With Cross-Task Class Discrimination in Online Continual Learning , Y. Guo, B. Liu, and D. Zhao. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 11878–11887

  67. [76]

    Episodic and Semantic Memory , E. Tulving. In Organization of Memory, Academic Press, 1972

  68. [77]

    Han and J

    Selecting Related Knowledge via Efficient Channel Attention for Online Continual Learning , Y. Han and J. Liu. In Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), 2022, pp. 1–7. [Online]. Available: https://ieeexplore.ieee.org/document/9742207

  69. [78]

    Supervised Contrastive Replay: Revisiting the Nearest Class Mean Classifier in Online Class-Incremental Continual Learning, Z. Mai, R. Li, H. Kim, and S. Sanner. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 3589–3599

  70. [79]

    Random Sampling with a Reservoir , J. S. Vitter. ACM Transactions on Mathematical Software (TOMS), vol. 11, no. 1, pp. 37–57, 1985. [Online]. Available: https://dl.acm.org/doi/10.1145/3145.3165

  71. [80]

    Chaudhry, P

    Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence , A. Chaudhry, P. K. Dokania, T. Ajanthan, and P. H. S. Torr. In Computer Vision – ECCV 2018, Springer International Publishing, 2018, pp. 556–572. [Online]. Available: http://dx.doi.org/10.10...

  72. [81]

    Continual Normalization: Rethinking Batch Normalization for Online Continual Learning , Q. Pham, C. Liu, and S. Hoi. arXiv preprint arXiv:2203.16102, 2022. [Online]. Available: https://arxiv.org/abs/2203.16102

  73. [82]

    Han and J

    Online Continual Learning via the Knowledge Invariant and Spread-Out Properties , Y. Han and J. Liu. Expert Systems with Applications, vol. 213, 2023, p. 119004. [Online]. Available: http://dx.doi.org/10.1016/j.eswa.2022.119004

  74. [83]

    An online continual object detector on VHR remote sensing images with class imbalance , X. Chen, J. Jiang, Z. Li, H. Qi, Q. Li, J. Liu, L. Zheng, M. Liu, and Y. Deng. Engi- neering Applications of Artificial Intelligence, vol. 117, 2023, p. 105549. [Online]. Available: https:/...

  75. [84]

    Jiang, Z

    Multi-instance Reservoir Sampling and Selection for Online Continual Detection over VHR Remote Sensing Images , J. Jiang, Z. Han, S. Wang, and C. Wang. In 2021 14th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), 2021,...

  76. [85]

    Kwon and T

    Toward an online continual learning architecture for intrusion detection of video surveillance , B. Kwon and T. Kim. IEEE Access, vol. 10, 2022, pp. 89732–89744. [Online]. Available: IEEE

  77. [86]

    Lopez-Paz and M

    Gradient Episodic Memory for Continual Learning , D. Lopez-Paz and M. Ranzato. arXiv preprint arXiv:1706.08840, 2022. [Online]. Available: https://arxiv.org/abs/1706.08840

  78. [87]

    Trust-region adaptive frequency for online continual learning , Y. Kong, L. Liu, M. Qiao, Z. Wang, and D. Tao. International Journal of Computer Vision, vol. 131, no. 7, 2023, pp. 1825–1839. Springer

  79. [88]

    Online continual learning on sequences , G. I. Parisi and V. Lomonaco. In Recent Trends in Learning From Data: Tutorials from the INNS Big Data and Deep Learning Conference (INNSBDDL2019), 2020, pp. 197–221. Springer

  80. [89]

    Online continual learning for embedded devices , T. L. Hayes and C. Kanan. arXiv preprint arXiv:2203.10681, 2022

  81. [90]

    When meta-learning meets online and continual learning: A survey , J. Son, S. Lee, and G. Kim. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. IEEE

  82. [91]

    Soutif-Cormerais, A

    A comprehensive empirical evaluation on online continual learning , A. Soutif-Cormerais, A. Carta, A. Cossu, J. Hurtado, V. Lomonaco, J. Van de Weijer, and H. Hemati. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3518–3528. 34

  83. [92]

    Reducing the dimensionality of data with neural networks , G. E. Hinton and R. R. Salakhutdinov. Science, vol. 313, no. 5786, 2006, pp. 504–507. American Association for the Advancement of Science

  84. [93]

    Attention is all you need , A. Vaswani. In Advances in Neural Information Processing Systems, 2017

  85. [94]

    Goodfellow

    Deep learning, I. Goodfellow. MIT Press, 2016

  86. [95]

    Large-scale machine learning with stochastic gradient descent , L. Bottou. In Proceedings of COMP- STAT’2010: 19th International Conference on Computational Statistics, Paris, France, 2010, pp. 177–186. Springer

  87. [96]

    Adam: A method for stochastic optimization , D. P. Kingma. arXiv preprint arXiv:1412.6980, 2014

  88. [97]

    Zhang, J

    Lookahead optimizer: k steps forward, 1 step back , M. Zhang, J. Lucas, J. Ba, and G. E. Hinton. In Advances in Neural Information Processing Systems, vol. 32, 2019

  89. [98]

    Parikh and S

    Proximal algorithms, N. Parikh and S. Boyd. Foundations and Trends ® in Optimization, vol. 1, no. 3, 2014, pp. 127–239. Now Publishers, Inc

  90. [99]

    Colson, P

    An overview of bilevel optimization , B. Colson, P. Marcotte, and G. Savard. Annals of Operations Research, vol. 153, 2007, pp. 235–256. Springer

  91. [100]

    Numerical optimization, S. J. Wright. 2006

  92. [101]

    The cross-entropy method: A unified approach to combinatorial optimization, Monte-Carlo simulation, and machine learning , R. Y. Rubinstein and D. P. Kroese. Springer, vol. 133, 2004

  93. [102]

    Shorten and T

    A survey on image data augmentation for deep learning , C. Shorten and T. M. Khoshgoftaar. Journal of Big Data, vol. 6, no. 1, 2019, pp. 1–48. Springer

  94. [103]

    Cutmix: Regularization strategy to train strong classifiers with localizable features , S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6023–6032

  95. [104]

    Mensink, J

    Distance-based image classification: Generalizing to new classes at near-zero cost , T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 11, 2013, pp. 2624–2637. IEEE

  96. [105]

    Paszke, S

    Pytorch: An imperative style, high-performance deep learning library , A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, and others. In Advances in Neural Information Processing Systems, vol. 32, 2019

  97. [106]

    Lomonaco, L

    Avalanche: an end-to-end library for continual learning , V. Lomonaco, L. Pellegrini, A. Cossu, A. Carta, G. Graffieti, T. L. Hayes, M. De Lange, M. Masana, J. Pomponi, G. M. Van de Ven, and others. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...

  98. [107]

    Pourcel, N.-S

    Online task-free continual learning with dynamic sparse distributed memory , J. Pourcel, N.-S. Vu, and R. M. French. In European Conference on Computer Vision, 2022, pp. 739–756

  99. [108]

    Aljundi, M

    Gradient-based sample selection for online continual learning , R. Aljundi, M. Lin, B. Goujaud, and Y. Bengio. In Advances in Neural Information Processing Systems, vol. 32, 2019

  100. [109]

    Sangermano, A

    Sample condensation in online continual learning , M. Sangermano, A. Carta, A. Cossu, and D. Bacciu. In 2022 International Joint Conference on Neural Networks (IJCNN), 2022, pp. 01–08. IEEE

  101. [110]

    Online continual learning in image classification: An empirical survey , Z. Mai, R. Li, J. Jeong, D. Quispe, H. Kim, and S. Sanner. Neurocomputing, vol. 469, pp. 28–51, 2022. Elsevier. 35

  102. [111]

    Online continual learning using enhanced random vector functional link networks , C. S. Y. Wong, G. Yang, A. Ambikapathi, and R. Savitha. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 1905–1909

  103. [112]

    Three scenarios for continual learning , G. M. Van de Ven and A. S. Tolias. arXiv preprint arXiv:1904.07734, 2019

  104. [113]

    Hadsell, D

    Embracing change: Continual learning in deep neural networks , R. Hadsell, D. Rao, A. A. Rusu, and R. Pascanu. Trends in Cognitive Sciences, vol. 24, no. 12, pp. 1028–1040, 2020. Elsevier

  105. [114]

    Han and J

    Online continual learning via the meta-learning update with multi-scale knowledge distillation and data augmentation, Y. Han and J. Liu. Engineering Applications of Artificial Intelligence, vol. 113, 2022, p. 104966. [Online]. Available: https://doi.org/10.1016/j.engappai.2022.104966

  106. [115]

    Continual lifelong learning with neural networks: A review , G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter. Neural Networks, vol. 113, pp. 54–71, 2019. [Online]. Available: https://doi.org/10.1016/j.neunet.2019.01.002

  107. [116]

    Lifelong learning with dynamically expandable networks , J. Yoon, E. Yang, J. Lee, and S. J. Hwang. arXiv preprint arXiv:1708.01547, 2017

  108. [117]

    Progressive neural networks, A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell. arXiv preprint arXiv:1606.04671, 2016

  109. [118]

    Wiewel and B

    Entropy-based sample selection for online continual learning , F. Wiewel and B. Yang. In 2020 28th European Signal Processing Conference (EUSIPCO), 2021, pp. 1477–1481. [Online]. Available: https://doi.org/10.23919/EUSIPCO49076.2021.9412831

  110. [119]

    Online continual learning under extreme memory constraints , E. Fini, S. Lathuiliere, E. Sangineto, M. Nabi, and E. Ricci. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII, Springer, 2020, pp. 720–735. [Online]. ...

  111. [120]

    Zhang, B

    A simple but strong baseline for online continual learning: Repeated augmented rehearsal , Y. Zhang, B. Pfahringer, E. Frank, A. Bifet, N. J. S. Lim, and Y. Jia. Advances in Neural Information Processing Systems, vol. 35, pp. 14771–14783, 2022

  112. [121]

    Summarizing Stream Data for Memory-Constrained Online Continual Learning , J. Gu, K. Wang, W. Jiang, and Y. You. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 2024, pp. 12217–12225. [Online]. Available: https://doi.org/10.1609/aaai.v38i1.11390

  113. [122]

    Rainbow memory: Continual learning with a memory of diverse samples , J. Bang, H. Kim, Y. Yoo, J. Ha, and J. Choi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8218–8227. [Online]. Available: https://doi.org/10.1109/CVPR46437....

  114. [123]

    Chaudhry, M

    Efficient lifelong learning with A-GEM , A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny. arXiv preprint arXiv:1812.00420, 2018

  115. [124]

    The impact of model size on catastrophic forgetting in Online Continual Learning , E. Lee. arXiv preprint arXiv:2407.00176, 2024

  116. [125]

    Kirkpatrick, R

    Overcoming catastrophic forgetting in neural networks , J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al. Proceedings of the National Academy of Sciences, vol. 114, no. 13, pp. 3521–3526...

  117. [126]

    Online continual learning on a contaminated data stream with blurry task boundaries , J. Bang, H. Koh, S. Park, H. Song, J.-W. Ha, and J. Choi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9275–9284. 36

  118. [127]

    Online continual learning with natural distribution shifts: An empirical study with visual data , Z. Cai, O. Sener, and V. Koltun. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 8281–8290

  119. [128]

    Zou and T

    Efficient Meta-Learning for Continual Learning with Taylor Expansion Approximation , X. Zou and T. Lin. In 2022 International Joint Conference on Neural Networks (IJCNN), 2022, pp. 1–8. [Online]. Available: IEEE

  120. [129]

    A survey of on-device machine learning: An algorithms and learning theory perspective , S. Dhar, J. Guo, J. Liu, S. Tripathi, U. Kurup, and M. Shah. ACM Transactions on Internet of Things, vol. 2, no. 3, pp. 1–49, 2021. [Online]. Available: ACM

  121. [130]

    SIESTA: Efficient online continual learning with sleep , M. Y. Harun, J. Gallardo, T. L. Hayes, R. Kemker, and C. Kanan. arXiv preprint arXiv:2303.10725, 2023

  122. [131]

    Tran, G.-N

    Hyperparameter optimization for improving recognition efficiency of an adaptive learning system , D.-P. Tran, G.-N. Nguyen, and V.-D. Hoang. IEEE Access, vol. 8, pp. 160569–160580, 2020. [Online]. Available: IEEE

  123. [132]

    Prabhu, Z

    Online continual learning without the storage constraint , A. Prabhu, Z. Cai, P. Dokania, P. Torr, V. Koltun, and O. Sener. arXiv preprint arXiv:2305.09253, 2023

  124. [133]

    Improving information retention in large scale online continual learning , Z. Cai, V. Koltun, and O. Sener. arXiv preprint arXiv:2210.06401, 2022

  125. [134]

    Bilevel continual learning , Q. Pham, D. Sahoo, C. Liu, and S. C. Hoi. arXiv preprint arXiv:2007.15553, 2020

  126. [135]

    Caccia, E

    Online Learned Continual Compression with Stacked Quantization Modules , L. Caccia, E. Belilovsky, M. Caccia, and J. Pineau. 2019

  127. [136]

    Schiemer, L

    Online continual learning for human activity recognition , M. Schiemer, L. Fang, S. Dobson, and J. Ye. Pervasive and Mobile Computing, vol. 93, p. 101817, 2023. [Online]. Available: Elsevier

  128. [137]

    Just Say the Name: Online Continual Learning with Category Names Only via Data Generation , M. Seo, D. Misra, S. Cho, M. Lee, and J. Choi. arXiv preprint arXiv:2403.10853, 2024

  129. [138]

    New insights for the stability-plasticity dilemma in online continual learning , D. Jung, D. Lee, S. Hong, H. Jang, H. Bae, and S. Yoon. arXiv preprint arXiv:2302.08741, 2023

  130. [139]

    Michel, G

    Learning Representations on the Unit Sphere: Investigating Angular Gaussian and Von Mises-Fisher Distributions for Online Continual Learning , N. Michel, G. Chierchia, R. Negrel, and J.-F. Bercher. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. ...

  131. [140]

    Online continual learning on class incremental blurry task configuration with anytime inference , H. Koh, D. Kim, J.-W. Ha, and J. Choi. arXiv preprint arXiv:2110.10031, 2021

  132. [141]

    Adaptive orthogonal projection for batch and online continual learning , Y. Guo, W. Hu, D. Zhao, and B. Liu. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 6, pp. 6783–6791, 2022

  133. [142]

    He and F

    Online continual learning for visual food classification , J. He and F. Zhu. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2337–2346, 2021

  134. [143]

    KJ and V

    Meta-consolidation for continual learning , J. KJ and V. N. Balasubramanian. In Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 14374–14386

  135. [144]

    Chen, A.-C

    Mitigating forgetting in online continual learning via instance-aware parameterization , H.-J. Chen, A.-C. Cheng, D.-C. Juan, W. Wei, and M. Sun. In Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 17466–17477. 37

  136. [145]

    Chaudhry, M

    On tiny episodic memories in continual learning , A. Chaudhry, M. Rohrbach, M. Elhoseiny, T. Ajanthan, P. K. Dokania, P. H. S. Torr, and M. Ranzato. arXiv preprint arXiv:1902.10486, 2019

  137. [146]

    Remind your neural network to prevent catastrophic forgetting , T. L. Hayes, K. Kafle, R. Shrestha, M. Acharya, and C. Kanan. In European Conference on Computer Vision, 2020, pp. 466–483

  138. [147]

    ACAE-REMIND for online continual learning with compressed feature replay , K. Wang, J. van de Weijer, and L. Herranz. In Pattern Recognition Letters, vol. 150, 2021, pp. 122–129

  139. [148]

    Aljundi, E

    Online continual learning with maximal interfered retrieval , R. Aljundi, E. Belilovsky, T. Tuytelaars, L. Charlin, M. Caccia, M. Lin, and L. Page-Caccia. In Advances in Neural Information Processing Systems, vol. 32, 2019

  140. [149]

    W´ ojcik, W

    Neural Architecture for Online Ensemble Continual Learning , M. W´ ojcik, W. Ko´ sciukiewicz, T. Kaj- danowicz, and A. Gonczarek. arXiv preprint arXiv:2211.14963, 2022

  141. [150]

    Iscen, J

    Memory-efficient incremental learning through feature adaptation , A. Iscen, J. Zhang, S. Lazebnik, and C. Schmid. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23–28, 2020, Proceedings, Part XVI 16, pp. 699–715

  142. [151]

    Online class-incremental continual learning with adversarial shapley value , D. Shim, Z. Mai, J. Jeong, S. Sanner, H. Kim, and J. Jang. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 11, 2021, pp. 9630–9638

  143. [152]

    Krizhevsky, G

    Learning multiple layers of features from tiny images , A. Krizhevsky, G. Hinton, and others. Toronto, ON, Canada, 2009

  144. [153]

    Zenke, B

    Continual learning through synaptic intelligence , F. Zenke, B. Poole, and S. Ganguli. In International Conference on Machine Learning, 2017, pp. 3987–3995

  145. [154]

    Imagenet: A large-scale hierarchical image database , J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255

  146. [155]

    Rousseau, P

    TED-LIUM: an Automatic Speech Recognition dedicated corpus , A. Rousseau, P. Del´ eglise, and Y. Esteve. In LREC, 2012, pp. 125–129

  147. [156]

    The Caltech-UCSD Birds-200-2011 dataset, C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. California Institute of Technology, 2011

  148. [157]

    Lomonaco and D

    Core50: a new dataset and benchmark for continuous object recognition , V. Lomonaco and D. Maltoni. In Conference on Robot Learning, 2017, pp. 17–26

  149. [158]

    Le and X

    Tiny Imagenet Visual Recognition Challenge , Y. Le and X. Yang. CS 231N, vol. 7, no. 7, 2015, pp. 3

  150. [159]

    Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms , H. Xiao, K. Rasul, and R. Vollgraf. arXiv preprint arXiv:1708.07747, 2017

  151. [160]

    Vinyals, C

    Matching networks for one shot learning , O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al. In Advances in Neural Information Processing Systems, vol. 29, 2016

  152. [161]

    Not just selection, but exploration: Online class-incremental continual learning via dual view consistency , Y. Gu, X. Yang, K. Wei, and C. Deng. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 7442–7451

  153. [162]

    He and F

    Online continual learning via candidates voting , J. He and F. Zhu. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 3154–3163

  154. [163]

    He and F

    Exemplar-free online continual learning , J. He and F. Zhu. In 2022 IEEE Interna- tional Conference on Image Processing (ICIP), 2022, pp. 541–545. [Online]. Available: https://ieeexplore.ieee.org/document/9769984 38

  155. [164]

    Scalable adversarial online continual learning, T. Dam, M. Pratama, M. D. M. Ferdaus, S. Anavatti, and H. Abbas. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2022, pp. 373–389. [Online]. Available: https://link.springer.com/chapter/10....

  156. [165]

    PCR: Proxy-based contrastive replay for online class-incremental continual learning , H. Lin, B. Zhang, S. Feng, X. Li, and Y. Ye. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 24246–24255

  157. [166]

    Online continual learning with declarative memory , Z. Xiao, Z. Du, R. Wang, R. Gan, and J. Li. Neural Networks, vol. 163, pp. 146–155, 2023. [Online]. Available: https://doi.org/10.1016/j.neunet.2023.01.011

  158. [167]

    Re-evaluating continual learning scenarios: A categorization and case for strong baselines , Y. Hsu. arXiv preprint arXiv:1810.12488, 2018

  159. [168]

    Aljundi, K

    Task-free continual learning, R. Aljundi, K. Kelchtermans, and T. Tuytelaars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 11254–11263

  160. [169]

    Buzzega, M

    Dark experience for general continual learning: A strong, simple baseline , P. Buzzega, M. Boschini, A. Porrello, D. Abati, and S. Calderara. In Advances in Neural Information Processing Systems, vol. 33, pp. 15920–15930, 2020

  161. [170]

    ERNIE 2.0: A continual pre-training framework for language understanding , Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 05, pp. 8968–8975, 2020. [Online]. Available: https://oai.org...

  162. [171]

    Online Continual Learning with Contrastive Vision Transformer , Z. Wang, L. Liu, Y. Kong, J. Guo, and D. Tao. In Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Ciss´ e, G. M. Farinella, T. Hassner, Eds. Springer Nature Switzerland, Cham, 2022, pp. 631–650

  163. [172]

    Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks , N. Reimers. arXiv preprint arXiv:1908.10084, 2019

  164. [173]

    Places: A 10 million Image Database for Scene Recognition , B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017. [Online]. Available: https://doi.org/10.1109/TPAMI.2017.2717679

  165. [174]

    Microsoft COCO: Common Objects in Context , T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ ar, and C. L. Zitnick. In Computer Vision – ECCV 2014, D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds. Springer International Publishing, Cham, 2014,...

  166. [175]

    Geiger, P

    Vision meets robotics: The kitti dataset , A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. The International Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013

  167. [176]

    Vip-deeplab: Learning visual perception with depth-aware video panoptic segmentation , S. Qiao, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3997–4008

  168. [177]

    Cordts, M

    The cityscapes dataset for semantic urban scene understanding , M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3213–3223

  169. [178]

    Maddern, G

    1 year, 1000 km: The oxford robotcar dataset , W. Maddern, G. Pascoe, C. Linegar, and P. Newman. The International Journal of Robotics Research, vol. 36, no. 1, pp. 3–15, 2017

  170. [179]

    DOTA: A large-scale dataset for object detection in aerial images , G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3974–3983. 39

  171. [180]

    Cheng, Z

    Structured Object-Level Relational Reasoning CNN-Based Target Detection Algorithm in a Remote Sensing Image , B. Cheng, Z. Li, B. Xu, X. Yao, Z. Ding, and T. Qin. Remote Sensing, vol. 13, no. 2, pp. 281, 2021. [Online]. Available: https://doi.org/10.3390/rs13020281

  172. [181]

    Zhang, T

    Remote Sensing Object Detection Meets Deep Learning: A metareview of challenges and advances , X. Zhang, T. Zhang, G. Wang, P. Zhu, X. Tang, X. Jia, and L. Jiao. IEEE Geoscience and Remote Sensing Mag- azine, vol. 11, no. 4, pp. 8–44, Dec. 2023. [Online]. Available: https://do...

  173. [182]

    Object detection in optical remote sensing images: A survey and a new benchmark , K. Li, G. Wan, G. Cheng, L. Meng, and J. Han. ISPRS Journal of Photogrammetry and Remote Sensing, vol. 159, pp. 296–307, 2020. [Online]. Available: https://doi.org/10.1016/j.isprsjprs.2019.11.008

  174. [183]

    Patino, T

    PETS 2017: Dataset and Challenge , L. Patino, T. Nawaz, T. Cane, and J. Ferryman. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPR W), 2017, pp. 2126–2132. [Online]. Available: https://doi.org/10.1109/CVPR W.2017.264

  175. [184]

    Thomee, D

    YFCC100M: The new data in multimedia research , B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li. Communications of the ACM, vol. 59, no. 2, pp. 64–73,

  176. [185]

    NUS-WIDE: A real-world web image database from National University of Singapore , T.-S. Chua, J. Tang, R. Hong, H. Li, Z. Luo, and Y. Zheng. In Proceedings of the ACM International Conference on Image and Video Retrieval, 2009, pp. 1–9. [Online]. Available: https://doi.org/10....

  177. [186]

    WebVision database: Visual learning and understanding from web data , W. Li, L. Wang, W. Li, E. Agustsson, and L. Van Gool. arXiv preprint arXiv:1708.02862, 2017. [Online]. Available: https://arxiv.org/abs/1708.02862

  178. [187]

    Reiss and D

    Introducing a New Benchmarked Dataset for Activity Monitoring , A. Reiss and D. Stricker. In Pro- ceedings of the 2012 16th International Symposium on Wearable Computers, 2012, pp. 108–109. [Online]. Available: https://doi.org/10.1109/ISWC.2012.13

  179. [188]

    ChatGPT: Language Model for Conversational AI , OpenAI. 2023. [Online]. Available: https://www.openai.com/chatgpt 40 A Newman’s Guideline

  180. [191]

    It is about creating focused questions to guide the study

    Develop Research Questions : This step is the base of the review. It is about creating focused questions to guide the study. These questions should be detailed enough to find relevant studies but also wide enough to cover all important areas. Clear research questions decide th...

  181. [192]

    It shows the main ideas, variables, and their connections

    Design Conceptual F ramework: Making a conceptual framework gives a structure for organizing and analyzing the literature. It shows the main ideas, variables, and their connections. This framework helps in choosing studies and ensures the analysis stays on track with the study...

  182. [193]

    These rules bring consistency to the process and remove bias, making sure only relevant and high-quality studies are part of the review

    Construct Selection Criteria : The selection criteria set rules for including or excluding studies. These rules bring consistency to the process and remove bias, making sure only relevant and high-quality studies are part of the review. This step makes the results more trustworthy

  183. [194]

    By using chosen keywords, Boolean operators, and clear criteria, this step ensures a complete collection of useful papers

    Develop Search Strategy: The search strategy is a plan for finding relevant studies from different databases. By using chosen keywords, Boolean operators, and clear criteria, this step ensures a complete collection of useful papers. A strong search strategy avoids missing impo...

  184. [195]

    It helps to ensure that only the studies directly connected to the research questions are included

    Select Studies Using Selection Criteria : In this step, we pick the studies that match the selection criteria. It helps to ensure that only the studies directly connected to the research questions are included. This step reduces confusion and keeps the review focused

  185. [196]

    This step makes the data easy to analyze and directly linked to the research questions

    Code Studies : Coding means organizing the important information from each study into specific categories. This step makes the data easy to analyze and directly linked to the research questions. It helps in finding insights in a structured way

  186. [197]

    This ensures the selected studies are good enough to provide reliable results

    Assess the Quality of Studies : After deciding which studies to include, we check their quality. This ensures the selected studies are good enough to provide reliable results. Poor-quality studies are excluded to keep the review rigorous

  187. [198]

    This involves finding patterns, similarities, or differences across studies and making conclusions based on these findings

    Synthesize Results of Individual Studies to Answer the Research Questions : In this step, we combine the results of the selected studies to address the research questions. This involves finding patterns, similarities, or differences across studies and making conclusions based ...

  188. [199]

    This includes summaries of the conclusions, study limitations, and suggestions for future research

    Report Findings : The last step is sharing the results in a clear and organized way. This includes summaries of the conclusions, study limitations, and suggestions for future research. This step provides valuable knowledge for the field. 41 B Components Tables Table 3: Mapping...

  189. [2016]

    Available: https://doi.org/10.1145/2894796.2894797

    [Online]. Available: https://doi.org/10.1145/2894796.2894797

  190. [2021]

    Available: https://doi.org/10.1109/ICCV.2021.10829

    [Online]. Available: https://doi.org/10.1109/ICCV.2021.10829

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.