{"id":"ffd59a05-f476-4151-8a3a-f24d6b50bbb9","arxiv_id":"2508.17455","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 58 multi-label data stream classifiers finds that concept drift is addressed by about half the papers, while label latency and concept evolution are each handled by only a small minority, and recurring labels are not handled at all.","lead":"This paper is a systematic review of 58 papers on multi-label data stream classification, organizing them into a hierarchy and comparing how they handle concept drift, label latency, concept evolution, and evaluation. It maps a fast-moving field and shows that label latency and new-label emergence are far less studied than drift, pointing to open problems for researchers and practitioners.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal count inconsistencies undermine the quantified gap analysis that is the review's central contribution.","rationale":"The review is well-structured, applies a recognized SLR framework, and provides a useful hierarchy and evaluation-metric discussion. However, the most load-bearing premise is not the search protocol's representativeness, as the reader emphasized, but the accuracy and consistency of the data extraction itself. The manuscript's internal count inconsistencies (58 vs 59 vs 52) appear in the very sections that produce the headline gap findings, and the citation error in Table 8 shows that the extraction step is not fully reliable. This directly threatens the central claim's quantitative conclusions about label latency and concept evolution being the least addressed problems. The complexity-table limitation is also real and self-admitted, but it is a weaker concern because the table can still be a useful compilation even if not rigorously verified. A clean re-code of the 58 papers would settle whether the gap counts hold. Because the reader already assigned CONDITIONAL and identified related issues, my read does not change the verdict: conditional acceptance remains appropriate until the extraction is audited and the counts corrected. Credit is due for the transparent methodology, the explicit listing of selection criteria, and the honest acknowledgment of the complexity table's limitations in Section 12.","tokens_in":39324,"tokens_out":3979,"duration_ms":39861,"concrete_test":"Independently re-code all 58 papers in Table 8 for the three binary variables Q2 (label latency), Q3 (concept drift handling), and Q4 (concept evolution handling), reading each paper's abstract and methods section, and compare the resulting counts with the 6, 32, and 8 reported in Sections 6, 7, and 8. Also recount the selected studies in Section 3.5 and verify the entries of Table 8 against the reference list, especially entry 15. If the corrected counts differ by more than one or two papers, or if the denominator is not consistently 58, the gap analysis must be revised and the abstract's quantified claims weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the review provides a systematic, comprehensive map of multi-label data stream classifiers, with quantified gaps: only 6 of 58 papers handle label latency, only 8 handle concept evolution, and none handle recurring labels. These counts are the empirical payoff of the SLR and the basis for the paper's main conclusions. Yet the paper's own core sections report inconsistent denominators: Section 7 says '32 out of the 59 studies' (Q3), Section 8 says 'of all the 52 investigated studies, only eight' (Q4), while Sections 3.5, 5.6, 6 and elsewhere consistently say 58 studies. Table 8 also miscites entry 15 as [90] when the title matches [93], indicating citation-level extraction errors. If the denominator is unstable, the binary coding of each paper (handles drift? handles evolution? handles latency?) is also suspect; even a few misclassifications would shift the headline gap percentages and the conclusion that label latency and concept evolution are the least addressed problems. Additionally, the 'exhaustive listing of asymptotic complexities' is explicitly self-limited: Section 12 admits Table 10 'lacks rigorous mathematical evaluation,' so the abstract's 'exhaustive' wording overstates the support for that contribution. The load-bearing assumption is that the extracted facts about 58 papers are accurate and consistently coded; the manuscript itself provides direct evidence this assumption is insecure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review (SLR) of multi-label data stream classification methods, following the Kitchenham and Charters framework. The authors describe a structured search across five databases, an inclusion/exclusion protocol, a quality assessment with a cutoff score of 3, and a final corpus of 58 studies. The core of the review is organized around six research questions: it proposes a hierarchy of methods (algorithm adaptation, problem transformation, ensemble), analyzes how methods handle label latency, concept drift, and concept evolution, discusses evaluation strategies and metrics, and identifies research gaps. The main claimed contributions are a full hierarchy of surveyed methods, an exhaustive listing of asymptotic complexities, and an identification of the least addressed problems, namely label latency and concept evolution, with no surveyed method handling recurring labels.","tokens_in":39550,"tokens_out":2929,"duration_ms":31601,"significance":"If the corpus and the extracted facts are accurate, this review would provide a valuable structured map of an active research area and quantitative evidence about under-studied problems such as label latency and concept evolution. The authors are transparent about their methodology, present a detailed data extraction form, and explicitly list limitations in Section 12, including the lack of rigorous verification of the complexity table. These are strengths for a survey paper. The main risk is that the quantified gap analysis rests on the consistency and correctness of the study counts and the binary coding of each paper's capabilities, and the manuscript currently contains internal inconsistencies that undermine this foundation. The contribution is therefore significant but needs correction before it can be fully relied upon.","major_comments":[{"comment":"The study count is inconsistent across the paper. Table 8 lists 58 studies and Section 3.5 states that 58 studies were selected, but Section 6 states 'of all the 58 investigated studies', Section 7 states '32 out of the 59 studies', and Section 8 states 'of all the 52 investigated studies, only eight'. These different denominators directly affect the headline quantitative claims that only 6 of 58 papers handle label latency and only 8 of 52 handle concept evolution. The authors must correct these counts, ensure all sections use the same final corpus size, and if any studies were excluded after data extraction, describe when and why this occurred.","section":"Sections 3.5, 6, 7, 8 and Table 8"},{"comment":"Entry 15 of Table 8 lists 'An Online Variational Inference and Ensemble Based Multi-label Classifier for Data Streams' as citation [90], but reference [90] is 'Multi-label classification via incremental clustering on an evolving data stream', while the cited title matches reference [93]. Entry 21 also cites [90] for 'Multi-label classification via incremental clustering on an evolving data stream'. This suggests a citation mapping error, and it raises the possibility that other entries in Table 8 have similar mismatches. The authors should verify every entry in Table 8 against its reference and correct any mis-associations, since accurate study identification is essential for a systematic review.","section":"Table 8, entries 15 and 21, with references [90] and [93]"},{"comment":"The abstract and the introduction's contribution list describe 'an exhaustive listing of the asymptotic complexities', but Section 12 explicitly acknowledges a 'lack of a rigorous mathematical evaluation of the asymptotic complexities in Table 10'. This is an overstatement of the support for that contribution. The authors should either temper the wording in the abstract and introduction to match the actual verification level or provide a rigorous verification and benchmarking of the entries in Table 10. In addition, Table 10 contains undefined symbols (e.g., 'p = ?' for PSLT) and unexpanded notation, which should be clarified.","section":"Section 4.5, Table 10, and Section 12"},{"comment":"The quality assessment procedure combines a citation-age score (Equation 1) with a four-item checklist, but the paper does not report the number of studies excluded at each stage of Figure 3 or the distribution of quality scores. Reporting these numbers would strengthen the reproducibility of the SLR and allow readers to assess whether the cutoff of 3 is appropriate. Without this information, the claim that the selected set is representative is harder to evaluate.","section":"Section 3.4 and Figure 3"}],"minor_comments":[{"comment":"The text states the number of publications per year 'rising in 2028', which is a typo and should read '2018'.","section":"Section 4.2, Figure 6"},{"comment":"The heading 'Co-occurence' is misspelled; it should be 'Co-occurrence'.","section":"Section 4.1 heading"},{"comment":"In the 'Is the code available' field, the listed possible values are 'Yes' and 'Not mentioned', with no 'No' option; this may not accurately capture papers that explicitly state code is not available.","section":"Table 9"},{"comment":"Several placeholders in the references and the ACM format header (e.g., 'https://doi.org/10.1145/nnnnnnn.nnnnnnn') remain unfilled; these should be completed or removed before publication.","section":"General presentation"}],"recommendation":"major_revision","confidential_remarks":"The inconsistent counts and the citation mismatch in Table 8 are concrete, fixable issues, but they strike at the quantitative gap analysis that forms the paper's main novelty. I would encourage the editor to require a full consistency check of the corpus and the extracted facts as part of the revision. The paper's honest disclosure of its own limitations in Section 12 is a positive sign, and the topic is within the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a genuine and mostly well-executed systematic review of multi-label data stream classification, covering work from 2016–2024. What it does best is synthesize a scattered field: it builds a clear hierarchy of methods (algorithm adaptation, problem transformation, ensembles), compiles evaluation practices, and gives a structured look at label latency, concept drift, and concept evolution. The focus on label latency is a real improvement over earlier surveys, and the gap analysis—showing that very few papers handle latency or concept evolution, and none handle recurring labels—is exactly the kind of map that helps people decide what to work on next. The SLR methodology is described in enough detail to be followed, which I appreciate.\n\nThe soft spots are real, and they are not minor. The paper gives inconsistent denominators: Section 6 says 58 studies, Section 7 says 59, Section 8 says 52. Those numbers are the basis for the headline conclusions, so this is a load-bearing inconsistency. Also, Table 8 entry 15 is cited as [90] but the title matches [93]—this suggests the study-level coding may be sloppy. And while the abstract promises an 'exhaustive listing' of asymptotic complexities, Section 12 concedes that Table 10 lacks rigorous mathematical evaluation; the wording overstates the support.\n\nTo be fair, the qualitative conclusions likely survive these issues: even if the exact counts shift by a few papers, the picture that drift is more studied than evolution or latency is probably right. So the flaws are serious but probably fixable. The authors have been honest about the limitations of the complexity table, which makes me more inclined to trust the rest.\n\nIf I were assigning this for peer review, I'd send it out with a clear request: reconcile the study counts, double-check the citation list and extraction consistency, and soften claims that the complexity listing is exhaustive. The reader who will get value here is a newcomer wanting a map of this subfield, or a researcher looking for gap statements to justify new work. It's worth engaging seriously, but only after a major revision.\n\nThe paper deserves a serious referee, not a desk rejection—just so the fixable problems are caught and the useful synthesis can be published in reliable form.","headline":"A useful but uneven SLR of multi-label data stream classification: the taxonomy and gap analysis are valuable, yet the paper's own inconsistent study counts and self-admitted limits on the complexity table undercut its most quantified claims.","tokens_in":40030,"tokens_out":1761,"would_cite":true,"duration_ms":20197,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 58 multi-label stream classification papers builds a method hierarchy and shows that label latency and concept evolution are the field's least-addressed problems.","keywords":["multi-label classification","data streams","concept drift","concept evolution","label latency","systematic literature review","asymptotic complexity","ensemble learning"],"falsifier":"Re-run the five-database query with additional sources, a relaxed quality cutoff, and the same 2016-2024 window, then count qualifying papers that explicitly handle finite delayed label latency or recurring classes; if the count rises substantially above six and eight (and above zero for recurring labels), the review's central gap findings would be overturned.","tokens_in":39169,"feed_emoji":"📊","tokens_out":5249,"duration_ms":51906,"temperature":0.7,"pith_summary":"This paper is a systematic literature review of multi-label classifiers for data streams. It surveys 58 papers published between 2016 and 2024, builds a full hierarchy of the methods by approach (algorithm adaptation, problem transformation, ensemble), lists their asymptotic complexities, and analyzes how they address concept drift, concept evolution, and label latency. The paper's central finding is that the field has concentrated on concept drift, while only six of 58 papers consider label latency and only eight consider concept evolution, with no paper handling recurring labels. A sympathetic reader would care because these are the problems that most affect real-world streams, where ground-truth labels arrive late and new classes appear without warning.","feed_headline":"Only 6 of 58 stream classifiers handle late labels, review finds","feed_subtitle":"Systematic review of multi-label data stream classifiers also finds no method that handles recurring new classes.","key_machinery":"The carrying object is the method hierarchy (Figure 9), which organizes all 58 surveyed methods into algorithm adaptation, problem transformation, and ensemble families and their subtypes (kNN, ELM, tree, SOM, powerset, binary relevance, regression-based, and others). The hierarchy does the work of making the literature comparable despite heterogeneous terminology, and the six research questions act as the extraction instrument that turns each paper into a row of comparable fields. The quality-assessment score and the asymptotic-complexity table are secondary instruments that let the review quantify which problems are addressed and at what cost.","core_discovery":"The authors claim to fill the gaps left by previous surveys by conducting a systematic review guided by six research questions covering how the classifier works, label latency, concept drift detection, concept evolution detection, evaluation strategy, and limitations. Their quantitative map of the literature shows that concept drift dominates attention (32 of 58 papers), while label latency is considered by only six papers, all addressing the infinite-latency case with none addressing feasible finite delay, and concept evolution by only eight, none of which handle recurring classes. The resulting hierarchy sorts the methods into algorithm adaptation, problem transformation, and ensemble approaches, and the complexity table records the asymptotic time and space behavior for 27 methods. The paper also reports that evaluation is far from standardized: only eleven papers are explicit about using prequential evaluation, and most papers are evaluated as a batch run at the end, making metric consensus an open problem.","pith_inferences":["If the corpus is representative, the next high-leverage contribution is a method that treats delayed (finite) label latency, since the review found no paper at all addressing that case.","The absence of recurring-label handling suggests a concrete testable design: a system that archives and reactivates label-specific models when an old class reappears, which the review implies but does not propose.","Because the review found only four semi-supervised methods, combining missing-label robustness with drift and evolution detection is likely to be a productive seam that the authors only signal implicitly.","The complexity claims should be treated as provisional: the authors state the table 'lacks rigorous mathematical evaluation,' so a benchmarking study that measures actual time and memory under identical conditions would be a natural follow-up the review does not itself perform."],"forward_implications":["Label latency is an open area: only six of 58 methods handle infinite label latency, and none handle finite delayed latency, so methods that work with delayed ground truth are a direct research opportunity.","Concept evolution is similarly under-addressed: only eight methods detect new labels and none handle recurring labels, so models that remember and re-introduce past classes are missing.","Concept drift dominates the field (32 of 58 methods), yet recurring drift is among the least handled drift patterns, so even within the most studied problem there is an open sub-problem.","Evaluation practice is not standardized: only eleven papers explicitly use prequential evaluation and F1 dominates while other metrics vary widely, which makes cross-method comparison unreliable.","The complexity table gives asymptotic bounds for 27 methods but the authors themselves note these lack rigorous mathematical evaluation, so the complexity map is a starting point rather than a verified benchmark."],"supporting_citations":[{"why":"Supplies the systematic-review methodology and structure the paper follows.","marker":"[70]"},{"why":"Supplies the PICOC criteria and the review process used to build the search and selection protocol.","marker":"[99]"},{"why":"The prior multi-label data-stream survey whose missing label-latency coverage and non-systematic approach motivate this review.","marker":"[145]"},{"why":"Listed as a 2022 survey with only partial coverage of multi-label stream classification, one of the gaps this review fills.","marker":"[57]"},{"why":"Listed as a 2021 review that covers novel class detection but not concept evolution and label latency fully, supporting the claimed gap.","marker":"[34]"}],"fun_headline_variants":["Only 6 of 58 stream classifiers handle label latency","No multi-label stream classifier addresses finite label delay","No method handles recurring new classes in multi-label streams","Review: label latency and recurring classes least addressed in stream ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's counts of who handles label latency or concept evolution are only as strong as its search-and-selection protocol, so a corpus that missed relevant papers would make those gaps look bigger than they are.","fun_headline_variants_meta":{"raw":{"variants":["Only 6 of 58 stream classifiers handle label latency","No multi-label stream classifier addresses finite label delay","No method handles recurring new classes in multi-label streams","Review: label latency and recurring classes least addressed in stream ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000353,"raw_usage":{"total_tokens":1873,"prompt_tokens":847,"completion_tokens":1026,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":961}},"tokens_in":463,"tokens_out":1026,"duration_ms":10303,"temperature":1.0,"reasoning_tokens":961,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:03:38.569390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the five-database query with additional sources, a relaxed quality cutoff, and the same 2016-2024 window, then count qualifying papers that explicitly handle finite delayed label latency or recurring classes; if the count rises substantially above six and eight (and above zero for recurring labels), the review's central gap findings would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior multi-label data-stream survey whose missing label-latency coverage and non-systematic approach motivate this review."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Listed as a 2021 review that covers novel class detection but not concept evolution and label latency fully, supporting the claimed gap."}],"review_version":1}