{"id":"a98de4d8-c07e-43b4-8c91-cd1d75f54100","arxiv_id":"2507.20614","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A systematic review proposes a five-part taxonomy for antisocial behavior prediction, covering early harm detection, harm emergence, propagation, behavioral risk, and proactive moderation.","lead":"This paper reviews 49 studies and organizes the emerging field of predicting online antisocial behavior into five task types. A generalist reader gets a structured map of how researchers forecast hate speech, derailment, and user risk before harm escalates.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.3's inclusion criteria are written in terms of the five taxonomy categories before the review, so the claimed empirical derivation of the taxonomy cannot be falsified by the 49-study corpus.","rationale":"The reader's weakest assumption identifies the same structural circularity I would flag: the taxonomy is introduced in Section 2.2 and then hard-coded into the inclusion criteria in Section 3.3, so the review cannot discover a predictive task type outside the five categories. This is not an accusation of bad faith; many surveys use a pre-specified coding scheme, and the paper could be read that way. But the text explicitly claims the taxonomy is 'derived' from the systematic review, and the abstract and contributions present it as a unifying framework for the field. That presentation makes the circularity load-bearing: if the taxonomy were presented as an analytic lens applied to the corpus, the circularity would be harmless, but as a derived result it is partly an artifact of the selection filter. The paper does have real strengths: a clear PRISMA-style protocol, explicit feature and dataset dimensions, an honest limitations section, and useful discussion of open challenges. The concern does not require rejection; it can be addressed by reframing the taxonomy as a coding scheme and by releasing the study list. The proposed test would settle whether the taxonomy is incomplete by asking whether taxonomy-blind coding of a broader retrieval set produces a category that does not fit. Until such a test is run, the conditional verdict is appropriate.","tokens_in":23045,"tokens_out":3163,"duration_ms":38607,"concrete_test":"Run an independent, taxonomy-blind mini-review: use the same databases and keyword strings from Section 3.1, but replace the Section 3.3 inclusion criteria with a generic criterion—'peer-reviewed ML study that predicts a future ASB-related outcome using classification or regression'—with no list of task types. Take a random sample of the resulting titles and abstracts, have two annotators independently open-code the predicted outcome for each, then map those codes onto the five taxonomy categories. If any code cannot be mapped (for example, predicting platform enforcement duration, predicting coordinated bot activity, or predicting user behavior after an unbanning decision that is not framed as reoffending), the taxonomy's completeness claim is undercut. Publish the full list of included studies and the coding decisions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that its five-category taxonomy organizes the space of ASB prediction tasks and is 'derived... grounded in a systematic review' (Section 2.2). The load-bearing premise is that the five categories emerged from, or at least are answerable to, the literature. That premise fails under the paper's own selection protocol: Section 3.3 includes only studies that 'explicitly aim to forecast one or more of the following: the early signals of harm, the emergence of antisocial behavior, its potential propagation, the behavioral risk posed by users, or outcomes relevant to proactive moderation.' These are exactly the five taxonomy categories. Exclusion reason (c) for the 32 full-text rejections is 'not addressing a task relevant to the taxonomy used in the study.' So every included study is guaranteed to fit one of the five labels; Figure 4's distribution is a consequence of the inclusion filter, not evidence about the structure of the field. A sixth task type would be excluded by construction. Section 8's admission that the author supplemented searches with a 'curated archive' makes independent verification harder, and no full study list is provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a systematic literature review of computational work that predicts antisocial behavior (ASB) on social media. Its central contribution is a proposed taxonomy of five task types—early harm detection, harm emergence prediction, harm propagation prediction, behavioral risk prediction, and proactive moderation support—organized along two dimensions (temporal orientation and operational purpose). The review reports a PRISMA-style selection process that yields 49 included studies, and it synthesizes the literature by task type, modeling technique, feature family, dataset structure, and platform. It also discusses open challenges such as linguistic generalization, cross-platform transfer, temporal drift, interpretability, and human-in-the-loop moderation. The paper claims that the taxonomy is derived from and grounded in the systematic review, and that the resulting map of the field can guide future research.","tokens_in":23310,"tokens_out":3718,"duration_ms":42222,"significance":"If the taxonomy and synthesis are accepted, the paper would provide a useful organizing framework for a fragmented research area, and its tables and figures give a compact overview of tasks, features, venues, and platforms. The manuscript has notable strengths: it follows a recognizable systematic-review structure with explicit research questions and a PRISMA-style flow diagram; it includes a candid limitations section; and it discusses ethical considerations that are often omitted from technical surveys. The feature-use table (Table 4) and platform-task observations are potentially useful references for researchers entering the area. However, the central claim that the taxonomy is empirically derived is undermined by the review's own inclusion criteria, which presuppose the five categories. Because the taxonomy is the paper's main contribution and the basis for its descriptive statistics, this issue is load-bearing. The corpus-level transparency is also insufficient for a systematic review: the full list of included studies and the per-study coding are not provided.","major_comments":[{"comment":"The inclusion criteria are written in terms of the five taxonomy categories before the review is conducted. Section 3.3 states that included studies must 'explicitly aim to forecast one or more of the following: the early signals of harm, the emergence of antisocial behavior, its potential propagation, the behavioral risk posed by users, or outcomes relevant to proactive moderation.' These are exactly the five categories introduced in Section 2.2, and exclusion reason (c) in Figure 1 rejects studies 'not addressing a task relevant to the taxonomy used in the study.' Consequently, every included study is guaranteed to fit at least one of the five labels, and the distribution in Figure 4 is a direct consequence of the inclusion filter rather than evidence about the structure of the field. This contradicts the claim in Section 2.2 that the taxonomy is 'derived... grounded in a systematic review,' and it means the review cannot discover a sixth task type by construction. The authors should either re-run the selection with open coding and report inter-annotator agreement, or explicitly reframe the taxonomy as a proposed analytic framework and describe the review as a mapping exercise rather than an empirical derivation.","section":"Section 3.3 vs. Section 2.2"},{"comment":"The review is not independently verifiable. The manuscript does not provide the list of 49 included studies, the extraction spreadsheet, or the category assignments for individual papers, and Section 3.4 indicates that screening and annotation were performed without reporting dual review or inter-annotator agreement. Section 8 additionally states that the author 'supplemented automated searches with a curated archive of domain-relevant publications,' which introduces a non-reproducible selection component. For a systematic review, the corpus and coding scheme should be released as supplementary material so that readers can check the taxonomy assignments and replicate the PRISMA counts. This is not a cosmetic concern: the paper's descriptive claims (e.g., Figure 4, Table 4, Table 5) depend entirely on which papers were included and how they were coded.","section":"Section 3.4 and Section 8"},{"comment":"The percentages in Figure 4 are not accompanied by a per-study mapping table, which makes it impossible to audit the category assignments. In addition, there are internal inconsistencies in the examples: Table 2 lists Hosseinmardi et al. [37] as Harm Emergence Prediction, while Section 5.1 illustrates ex-ante prediction with the Instagram cyberbullying example citing Hosseinmardi et al. [56]; reference [56] is the 2015 arXiv version of the same group's work. The authors should provide a supplementary table linking each included study to its taxonomy category, temporal strategy, granularity, feature families, and RQ2/RQ3 codes, and should reconcile duplicate and near-duplicate references.","section":"Section 5.1 and Figure 4"},{"comment":"The exponential growth claim is not statistically supported. Figure 2 plots observed counts against y = exp(0.205t) for t = 1,...,12, but the manuscript reports no goodness-of-fit measure, no confidence interval, no residual analysis, and no justification for choosing an exponential model over a descriptive bar chart. The text concludes 'a clear and accelerating growth trajectory,' but with yearly counts between 1 and 9 and an unusual dip in 2022, the fitted curve may not be a meaningful summary. The authors should either report standard fit diagnostics (e.g., R-squared, RMSE, model comparison) or present the raw counts descriptively without an overlaid fitted curve.","section":"Section 4.1, Figure 2"}],"minor_comments":[{"comment":"The sentence 'we derive a taxonomy grounded in a systematic review of recent literature (Section 4)' is misleading because Section 3.3 has already fixed the five categories as inclusion criteria; please rephrase to reflect the actual relationship between the taxonomy and the review.","section":"Section 2.2"},{"comment":"Google Scholar is not a reproducible search source because its ranking and coverage change over time; the authors should provide the exact query strings, search dates per database, and the version of any aggregator tool used, or drop Google Scholar from the database list.","section":"Section 3.2"},{"comment":"The Multiple Correspondence Analysis plot lacks details on preprocessing, the number of keywords considered, the threshold rationale, and the variance explained by the two displayed dimensions; without these, the cluster interpretation is difficult to evaluate.","section":"Section 4.3, Figure 3"},{"comment":"The reference list contains duplicates: [38] and [71] are the same paper, and [59] and [63] are the same paper; these should be deduplicated and cited consistently.","section":"References"},{"comment":"The claim that 'over half of the surveyed datasets exhibit conversational flow' is not backed by a count or by the table; please provide the underlying numbers or qualify the statement.","section":"Section 5.3"},{"comment":"The reference list in the ex-ante classification row includes a duplicated entry '[69, 69, 70]'; please clean up the citation list.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern raised by the reader is genuine and is visible in the manuscript itself: the inclusion criteria in Section 3.3 reproduce the taxonomy of Section 2.2 almost verbatim. The paper also lacks the usual transparency artifacts of a systematic review (full study list, extraction table, dual coding), and Section 8 openly acknowledges a non-reproducible 'curated archive' supplement. For this reason I would not accept the current version, but the core survey material is potentially useful. If the authors reframe the taxonomy as a proposed framework rather than an empirical discovery, and if they release the corpus-level data, a revision could become publishable. I would also ask the editor to ensure that the supplementary materials contain the per-paper coding table, since the descriptive statistics in Figures 4 and 5 and Tables 4 and 5 cannot otherwise be audited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, readable survey of a fragmented subfield, and the five-category taxonomy is a reasonable organizing device. But the central claim that the taxonomy is derived from the literature is circular: Section 3.3's inclusion criteria already require papers to forecast one of those five outcome types, so the review cannot discover anything outside them. That is a real limitation, not a dealbreaker.\n\nWhat is genuinely useful: the paper gives a clean map of ASB prediction tasks—early harm detection, harm emergence, harm propagation, behavioral risk, proactive moderation—and organizes the literature along prediction type, temporal strategy, granularity, features, models, and platforms. That is a service to a community that lacks a shared vocabulary. The PRISMA-style methodology is described in enough detail to be plausible. The limitations section is honest about coverage gaps, English-language bias, and the author's curated archive. The synthesis of open challenges (temporal drift, cross-platform transfer, benchmarking gaps) is sensible and appropriately cited.\n\nThe soft spots, in order. First, the circularity. The taxonomy is presented in Section 2.2 and then used as the screening instrument in Section 3.3; exclusion reason (c) is \"not addressing a task relevant to the taxonomy used in the study.\" So Figure 4's distribution is a function of the filter, not evidence about the field's structure. The paper would be stronger if it presented the taxonomy as an a priori coding scheme and explicitly acknowledged that the review cannot detect task types outside it. That is a conceptual fix, but it changes what the survey claims to show.\n\nSecond, reproducibility. There is no list of the 49 included papers, no extraction spreadsheet, and screening and coding appear to be single-author. For a systematic review, the study set is the data; without it, readers cannot check coverage or assignment. This is a moderate concern, and it interacts with the curated-archive supplement mentioned in Section 8.\n\nThe exponential fit in Figure 2 is just a descriptive curve; nothing hangs on it.\n\nBottom line: the contribution is organizational, not empirical. It deserves a serious referee who will ask for the study list and a reframing of the taxonomy as a coding scheme. I would send it to review—the subfield needs this kind of ordering—and the author is clearly thinking seriously about the work.","headline":"A competent, clearly written survey whose five-category taxonomy is genuinely useful, but the claim that the taxonomy is grounded in the review is undercut by the inclusion criteria being written in the taxonomy's own terms.","tokens_in":23806,"tokens_out":1909,"would_cite":false,"duration_ms":20614,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the fragmented research on predicting antisocial behavior online can be unified under a five-category taxonomy—early harm detection, harm emergence, harm propagation, behavioral risk, and proactive…","keywords":["antisocial behavior prediction","systematic literature review","hate speech forecasting","early harm detection","harm propagation prediction","behavioral risk prediction","proactive moderation","online toxicity forecasting"],"falsifier":"Re-run the literature search with inclusion criteria that do not mention the five task categories and count how many peer-reviewed machine-learning studies predicting future harmful outcomes, such as coordinated harassment campaigns, moderator workload, or platform policy violations, fall outside all five labels; even one such study would show the taxonomy is incomplete.","tokens_in":22777,"feed_emoji":"🛡️","tokens_out":6383,"duration_ms":67172,"temperature":0.7,"pith_summary":"ASB prediction, the computational modeling of future harmful behavior rather than detection of content already posted, is scattered across many task formulations. The paper sets out to unify this field with a taxonomy built on two dimensions, temporal orientation and operational purpose, yielding five task types: early harm detection, harm emergence prediction, harm propagation prediction, behavioral risk prediction, and proactive moderation support. It reviews 49 studies that fit its inclusion criteria and analyzes their tasks, features, models, and datasets, then lists open challenges such as English-only data, platform dependence, temporal drift, and missing benchmarks. A sympathetic reader would care because the taxonomy gives researchers a shared vocabulary for comparing methods and gives moderators a map of what can be forecast before harm unfolds.","feed_headline":"Review sorts 49 studies of online hate prediction into five tasks","feed_subtitle":"The field's five task types, from spotting harm early to proactive moderation, now have one shared map—and clear gaps.","key_machinery":"The load-bearing object is the taxonomy itself: a two-dimensional grid whose axes are temporal orientation (anticipating the emergence, escalation, or spread of harm) and operational purpose (supporting moderation, risk assessment, or intervention). The five categories are its cells, and the paper uses them as both the coding scheme for the literature review and the explanatory frame for comparing methods, features, and datasets. Secondary machinery includes the temporal-strategy distinction between ex-ante prediction and peeking strategies that use early interaction signals, along with the granularity levels (micro, meso, macro) adapted from information-cascade research.","core_discovery":"On the paper's own terms, the central discovery is that predictive ASB work, despite different labels and platforms, falls into five recurring task types distinguished by when in the harm lifecycle the prediction sits and what operational job it does. Early harm detection flags harm from the first few messages; harm emergence prediction asks whether a currently civil interaction will become toxic; harm propagation prediction estimates how far and fast harmful content will spread; behavioral risk prediction scores whether a user will reoffend, migrate to extreme communities, or become a target; proactive moderation support evaluates content before publication or ranks it for moderator review. The paper further reports that these categories are unevenly populated, with early detection and emergence dominating while proactive moderation is least studied, and that task formulation splits along classification versus regression, ex-ante versus peeking temporal strategies, and micro-, meso-, or macro-level granularity.","pith_inferences":["If the taxonomy is used as an inclusion filter rather than a descriptive result, it will systematically miss task types it did not predefine, such as forecasting coordinated inauthentic behavior, moderator workload, or platform-level policy outcomes; future reviews should derive categories inductively before fixing them.","The two axes suggest a testable separation: content-rich ex-ante moderation tasks may continue to favor classical feature-based models, while propagation and progressive-peeking tasks with temporal structure should benefit most from sequence and graph models; a benchmark organized by taxonomy category could confirm this.","Platform safety teams could use the reported distribution to prioritize investment: early detection already has many methods, whereas proactive moderation support is both least studied and most directly actionable for prevention.","One direct extension would be a shared task with one dataset per category and a uniform temporal split, which would turn the taxonomy from a descriptive map into an evaluation standard."],"forward_implications":["Researchers can position any new prediction task in one of five categories, enabling direct comparison of methods that were previously labeled inconsistently.","The reported distribution of effort (27.7% early detection, 23.4% emergence, 19.1% propagation, 17.0% behavioral risk, 12.8% proactive moderation) identifies proactive moderation support as the least developed category and a likely target for new work.","Task design is strongly shaped by platform structure: Twitter suits propagation and emergence tasks, while threaded Reddit and Wikipedia discussions suit derailment and early-detection tasks, implying that cross-platform generalization cannot be assumed.","The heavy concentration of English-language corpora (over 83%) and the lack of standardized benchmarks mean that multilingual modeling and shared evaluation tasks are the clearest levers for field-wide progress.","Temporal framing matters for accuracy: peeking and progressive prediction strategies that use early interaction signals are gaining ground for multi-turn and cascade tasks, and temporal drift is an acknowledged source of performance decay."],"supporting_citations":[{"why":"Foundational large-scale analysis showing antisocial user behavior is detectable before bans, grounding the behavioral risk prediction category.","marker":"[34]"},{"why":"Shows cyberbullying on Instagram can be predicted from pre-comment signals, grounding harm emergence prediction.","marker":"[37]"},{"why":"Demonstrates conversation structure predicts future toxicity across 58 million tweets, grounding early harm detection and structural features.","marker":"[10]"},{"why":"Supplies the systematic review procedure adapted for the paper's study selection.","marker":"[35]"},{"why":"Supplies the reporting flow that governs how the 49-study corpus is screened and selected.","marker":"[36]"},{"why":"Provides the thematic-analysis dimensions (prediction type, temporal strategy, granularity) used to characterize tasks.","marker":"[33]"},{"why":"Provides a continuous hate-severity regression model that illustrates behavioral risk prediction and is compared against an external toxicity scoring service.","marker":"[38]"}],"fun_headline_variants":["Five tasks unify online antisocial behavior prediction research","Predicting hate before it erupts: five tasks mapped from 49 studies","Online hate prediction: one taxonomy, five tasks, many gaps","Systematic review defines five task types for predicting online abuse","From early warning to proactive moderation: five ways to predict online harm"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy's five categories are fixed in advance by the review's inclusion criteria, so the review can only find studies that fit one of the five labels and cannot discover a genuinely new predictive task type.","fun_headline_variants_meta":{"raw":{"variants":["Five tasks unify online antisocial behavior prediction research","Predicting hate before it erupts: five tasks mapped from 49 studies","Online hate prediction: one taxonomy, five tasks, many gaps","Systematic review defines five task types for predicting online abuse","From early warning to proactive moderation: five ways to predict online harm"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000724,"raw_usage":{"total_tokens":3249,"prompt_tokens":948,"completion_tokens":2301,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":2214}},"tokens_in":564,"tokens_out":2301,"duration_ms":15953,"temperature":1.0,"reasoning_tokens":2214,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:39:26.246717+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the literature search with inclusion criteria that do not mention the five task categories and count how many peer-reviewed machine-learning studies predicting future harmful outcomes, such as coordinated harassment campaigns, moderator workload, or platform policy violations, fall outside all five labels; even one such study would show the taxonomy is incomplete.","supporting_citations":[{"cited_title":"In: Cha, M., Mascolo, C., Sandvig, C","cited_arxiv_id":null,"evidence_quote":"Foundational large-scale analysis showing antisocial user behavior is detectable before bans, grounding the behavioral risk prediction category."},{"cited_title":"In: Kumar, R., Caverlee, J., Tong, H","cited_arxiv_id":null,"evidence_quote":"Shows cyberbullying on Instagram can be predicted from pre-comment signals, grounding harm emergence prediction."},{"cited_title":"Technical report, Keele University, Department of Computer Science (2004)","cited_arxiv_id":null,"evidence_quote":"Supplies the systematic review procedure adapted for the paper's study selection."},{"cited_title":"Bmj 339 (2009)","cited_arxiv_id":null,"evidence_quote":"Supplies the reporting flow that governs how the 49-study corpus is screened and selected."},{"cited_title":"ACM Comput","cited_arxiv_id":null,"evidence_quote":"Provides the thematic-analysis dimensions (prediction type, temporal strategy, granularity) used to characterize tasks."}],"review_version":2}