{"id":"32809cfb-0588-4a26-a0c2-81cee0787930","arxiv_id":"2501.09309","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review of ML, DL, and NLP methods for detecting suicidal ideation from social media, summarizing prior studies and challenges without producing new experimental results.","lead":"This paper reviews how machine learning, deep learning, and natural language processing are used to detect suicidal thoughts in social media posts. It is a starting map for researchers, but it does not run new experiments or provide code or data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Conclusion's precision/recall range is not supported by the reviewed evidence: Table I contains values outside the stated bounds, and no derivation or selection criteria is given.","rationale":"The paper is a narrative review, not a new empirical study. The strongest claim is the aggregate precision and recall range in the conclusion. That claim is not supported by the paper's own Table I: the table contains precision and recall values outside the stated ranges, mixes metrics that are not interchangeable, and provides no methodology for aggregating them. The reader's weakest assumption identified the same underlying problem: that the surveyed datasets and metrics are comparable and that the reported performance transfers to a single headline range. My concern is more specific: even within the paper's own table, the stated range is not the range of the reported precision and recall values, so the problem is not merely external generalizability but internal support. This does not change the appropriate verdict. The work is still best classified as UNVERDICTED because it is a synthesis artifact without an original, testable claim. The correct remedy is not to reject the entire review but to remove or substantiate the unsupported aggregate range, which would require either a systematic review protocol or a clearly defined subset of studies. I therefore keep the reader's UNVERDICTED verdict unchanged.","tokens_in":24660,"tokens_out":2674,"duration_ms":34413,"concrete_test":"Reconstruct a per-study table from Table I extracting, for every row that reports precision and recall, the exact metric values and the model used. Then recompute the min-max ranges across all such rows. If the result is not 82-97% precision and 71-94% recall (e.g., because Kholifah's 50.24% precision or Rabani's 98.2% recall are included), then the conclusion's ranges must be revised to match the data or explicitly restricted to a justified subset of studies. A second check: if a subset is claimed, require the authors to list the studies included and the metric definition used for each.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in Section IV, is that 'Key findings demonstrate precision rates of 82-97% and recall rates of 71-94%' for SVM, Random Forest, and neural network models. This range is load-bearing because the abstract and conclusion use it to argue that these technologies have 'life-saving potential.' However, the claim is not derivable from Table I, which is the only quantitative synthesis provided. Table I reports heterogeneous metrics: accuracy, F1, AUC, precision, recall, and sensitivity, across different platforms, tasks (depression, suicidal ideation, suicide attempts), and evaluation protocols. Several entries directly contradict the stated ranges: Kholifah et al. [28] reports precision of 50.24% and recall of 70.89%; Rabani et al. [16] reports recall of 98.2%; Fodeh et al. [19] reports sensitivity of 0.912; Zhang et al. [21] reports recall of 94.9%. If the range is meant to summarize Table I, it omits these values; if it summarizes a subset, the subset and inclusion criteria are not stated. No meta-analytic method, protocol, or extraction rule is provided, so the range cannot be reproduced or checked. Because the headline performance claim rests entirely on this unsupported aggregate, the conclusion overstates what the reviewed literature shows.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative review of machine learning, deep learning, and natural language processing approaches for detecting suicidal ideation from social media content. It surveys roughly twenty studies, summarizes them in a comparison table (Table I), describes classical and deep learning models with equations, and concludes that such technologies achieve precision rates of 82–97% and recall rates of 71–94%, framing them as life-saving tools if developed responsibly. The paper also discusses challenges such as dataset bias, model interpretability, privacy, and ethical deployment.","tokens_in":24888,"tokens_out":1878,"duration_ms":20226,"significance":"If the surveyed evidence were synthesized rigorously, this review would address an important and timely interdisciplinary question: whether automated text classification can support suicide prevention. The paper collects a broad range of relevant studies and provides a useful taxonomy of methods, including lexicon-based tools, classical classifiers, neural architectures, and transformer-based models. However, the paper does not perform a systematic search, does not assess study quality, and aggregates heterogeneous metrics into a single performance range without justification. Its main quantitative conclusion is therefore not reproducible from the evidence presented. The topic is significant, but the manuscript in its current form does not deliver a defensible synthesis.","major_comments":[{"comment":"The claim that 'Key findings demonstrate precision rates of 82-97% and recall rates of 71-94% using models such as SVMs, Random Forests, and neural networks' is not supported by Table I, the only quantitative synthesis in the paper. Several entries in Table I fall outside these bounds: Kholifah et al. [28] reports precision of 50.24% and recall of 70.89%; Rabani et al. [16] reports recall of 98.2%; Fodeh et al. [19] reports sensitivity of 0.912; Zhang et al. [21] reports recall of 94.9%. If the range is intended to summarize a subset of studies, the subset and inclusion criteria are not stated. Without a stated extraction rule or meta-analytic method, the range cannot be reproduced or checked, and the conclusion overstates what the reviewed literature shows. This is a load-bearing issue because the abstract and conclusion use this range to argue for the technology's life-saving potential.","section":"Section IV (Conclusion)"},{"comment":"The review lacks a systematic methodology: there is no description of search databases, search strings, inclusion/exclusion criteria, or quality appraisal. The narrative in Section II is organized by study rather than by evidence level, and Table I mixes studies with different platforms (Twitter, Reddit, Facebook, KNHANES survey data), different tasks (suicidal ideation detection, suicide attempt prediction, depression detection, suicide note identification), and different evaluation metrics (accuracy, F1, AUC, precision, recall, sensitivity). Because these metrics measure different quantities under different class distributions and evaluation protocols, combining them into aggregate ranges is methodologically inappropriate. The manuscript needs at least a clear statement of which studies were included in any quantitative claim and a justification for why their metrics are comparable.","section":"Section II (Previous Works) and Table I"},{"comment":"Table I reports 'Results' as a mixture of accuracy, F1, AUC, precision, recall, and sensitivity, often with no indication of which class or threshold is being reported (e.g., 'Avg rate: 79%-87%' for Parraga-Alva et al. [14]; 'Strong performance in classification' for Haque et al. [22]). The text in Section III.F lists definitions of metrics but does not apply them consistently to the surveyed studies. This makes it impossible for a reader to determine whether a reported value is a precision, recall, or F1 score, or whether it is a macro-averaged versus micro-averaged figure. The authors should standardize the reported metrics and state extraction rules, or explicitly refrain from reporting aggregate ranges.","section":"Table I and Section III.F"}],"minor_comments":[{"comment":"The title uses 'It's Effect' in the header on page 342, which should be 'Its Effect'.","section":"Title"},{"comment":"The heading 'Evaluation matrices' should be 'Evaluation metrics'. This appears in the text before reference [88].","section":"Section III.F.4"},{"comment":"Table II provides a qualitative comparison of traditional ML, deep learning, and NLP, but it does not include any concrete performance values or citations for the claims about accuracy, interpretability, or suitability. Readers cannot verify the comparative claims from this table alone.","section":"Table II"},{"comment":"Figure 1 presents a useful taxonomy, but it is never referenced in the body text. Please add a sentence that directs the reader to Figure 1 and explains how the taxonomy maps to the sections in III.F.","section":"Figure 1"},{"comment":"References are formatted inconsistently (e.g., some entries include 'doi: 10.1007/s11920-018-0914-y.' as part of the publisher name, and several arXiv or conference papers lack page ranges or venue names). A consistent style per the journal's guidelines is needed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's core message—that ML/DL methods are being actively and often successfully applied to suicide ideation detection from social media—is broadly consistent with the literature. However, the headline quantitative claim in the conclusion is not supported by the paper's own Table I, and the lack of a systematic protocol means the review is not a reliable synthesis. I would like the authors to either substantially revise the quantitative claims to match the evidence (e.g., reporting ranges only for studies that report a given metric on a defined task, with explicit inclusion criteria) or add a carefully documented meta-analytic section. Given the sensitivity of the application domain, overstatement of performance ranges could mislead readers about readiness for real-world screening. The paper could be acceptable after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a narrative review, not a research contribution. It compiles known ML/DL/NLP approaches for detecting suicidal ideation in social media, gives a taxonomy, and discusses ethics. That is useful for newcomers. But the paper's headline quantitative claim — precision 82–97%, recall 71–94% — does not survive contact with its own table. Several entries fall outside those bounds (e.g., Kholifah's precision 50.24%, Rabani's recall 98.2%, Zhang's recall 94.9%). No subset or inclusion rule is given. That makes the central 'life-saving potential' argument rest on an unreproducible summary.\n\nWhat it does well: the taxonomy in Figure 1 is a readable map of the field; Table I, despite its problems, gives a sense of the range of datasets and platforms; the discussion of challenges (generalizability, privacy, black-box models, data imbalance) is sane. The survey covers recent work, including transformer-based methods, and correctly identifies the main open problems.\n\nWhere it is soft: beyond the precision/recall issue, the review is not systematic. No search strategy, no inclusion/exclusion criteria, no quality assessment. The studies in Table I differ wildly in task (depression vs. suicidal ideation vs. suicide attempt), platform, label definition, and evaluation metric; aggregating them into a single performance range is pointless. There are also smaller signs of haste: the title uses 'It's' instead of 'Its,' one author is 'Tianlin Zhag' in Table I, and the section on evaluation metrics is shallow. The paper's own earlier works are cited as background; that's fine.\n\nIs the central message wrong? No. The broad claim that ML/DL can detect suicidal ideation in social media at useful levels is supported by the cited literature. But the specific numbers are not. A careful revision that either reports the range of values from Table I with the caveat that they are not comparable, or drops the aggregate metric claim entirely, would fix the main problem.\n\nWho is this for: a student or practitioner looking for a quick entry point to the literature. Not for someone seeking a rigorous synthesis.\n\nRecommendation: I would not cite it in its current form. If a desk editor wants to give it a chance, send it to a referee with instructions to check the quantitative claims; the paper could be made sound with one round of major revision. As it stands, I'd treat it as a flawed but salvageable review.","headline":"A serviceable narrative review undercut by an unsupported precision/recall range that contradicts its own Table I.","tokens_in":25402,"tokens_out":2893,"would_cite":false,"duration_ms":59835,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of prior studies finds that machine learning on social media text can identify suicidal ideation with 82–97% precision and 71–94% recall, while noting that real-world deployment still faces data, interpretability, and ethical…","keywords":["suicidal ideation detection","social media analysis","mental health","text analysis","machine learning","deep learning","natural language processing","suicide prevention"],"falsifier":"Train one of the reviewed models on posts from one platform and time period, then test it on posts from a different platform and time period; if precision and recall fall well below the quoted 82–97% and 71–94% ranges, the transferable-screening claim collapses.","tokens_in":24481,"feed_emoji":"🧠","tokens_out":7844,"duration_ms":73142,"temperature":0.7,"pith_summary":"This review aims to establish that machine learning, deep learning, and natural language processing applied to social media text can detect signs of suicidal ideation. Drawing on a comparison of prior studies, it reports precision rates of 82–97% and recall rates of 71–94% for models such as SVMs, Random Forests, and neural networks. The paper argues these results make large-scale automated screening technically plausible, while insisting that real deployment still faces limited datasets, poor generalizability, opacity, and unresolved ethical questions.","feed_headline":"Social media text flags suicide risk at 82–97% precision","feed_subtitle":"The lab results are good, but bias, privacy, and interpretability still block real-world deployment.","key_machinery":"The load-bearing object is the comparison table of prior studies, which supplies the performance ranges quoted in the conclusion. Around it the paper builds a taxonomy of methods—lexicon-based emotion analysis, supervised and unsupervised machine learning, deep architectures, feature engineering, and evaluation metrics—and a standard pipeline of data collection, preprocessing, feature extraction, model training, and testing. The taxonomy explains how the reported results are generated, and the table is what lets the paper generalize across platforms and studies.","core_discovery":"The paper's central claim is that social media posts carry detectable linguistic traces of suicidal ideation, and that modern classifiers—SVMs, Random Forests, CNNs, LSTMs, and transformer models—can separate those traces from ordinary expression with reported precision of 82–97% and recall of 71–94%. It argues that deep learning improves context-aware detection over simpler lexicon and frequency-based approaches, because networks preserve word order and longer context. The same synthesis supports a second claim: this capability is not yet deployment-ready, since datasets are small, English-centric, and often biased, models are hard to interpret, and privacy and duty-of-care questions remain unresolved.","pith_inferences":["The quoted performance range likely overstates real-world value, because most of the underlying labels come from forum self-disclosure or platform flags rather than clinical assessment.","A useful benchmark the paper does not report would be out-of-distribution testing—training on one platform and time period and testing on another; this would directly measure generalizability.","The same text-analysis pipeline could extend to other crisis states such as depression, PTSD, or substance use, which share overlapping linguistic markers with suicidal ideation.","The hardest unresolved distinction is between expressing distress and declaring intent; future datasets should label near-term risk separately from general ideation, a distinction none of the surveyed studies cleanly makes."],"forward_implications":["If the reported precision and recall hold outside the original datasets, social media platforms could screen posts and route flagged accounts to human review.","Deep learning models such as LSTMs and transformers would be favored over keyword lexicons for detecting context-dependent distress.","Models would need careful handling of class imbalance, since suicidal posts are rare relative to ordinary content.","Any real deployment would require consent, privacy, and crisis-response protocols to avoid stigmatizing flagged users.","Because most training data is English and US-centric, the method must be revalidated in other languages and demographic groups before broad use."],"supporting_citations":[{"why":"Supplies the high-accuracy Twitter Random Forest result (98.7% precision, 98.2% recall) that anchors the quoted performance range.","marker":"[16]"},{"why":"Supplies the Reddit LSTM-CNN deep learning result (93.8% accuracy) used to support deep-learning detection claims.","marker":"[7]"},{"why":"Supplies SVM and decision-tree results on Twitter mental-health prediction, cited for very high accuracy figures.","marker":"[20]"},{"why":"Supplies large-scale Facebook data and ANN models whose AUC results support real-user post detection.","marker":"[23]"},{"why":"Supplies BERT and Sentence-BERT classification results that ground the paper's transformer-model claims.","marker":"[22]"},{"why":"Carries the paper's caution about practical and ethical hurdles to deploying machine learning for suicide prevention.","marker":"[5]"}],"fun_headline_variants":["Social posts betray suicide risk with 82–97% precision","AI reads suicide risk in social text—but ethics lag behind","Social media words detect suicidal thoughts, but bias remains","82–97% precision: social posts predict suicide risk, not yet deployable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the studies it combines use comparable datasets and metrics, so the quoted performance range predicts how a real screening tool would behave.","fun_headline_variants_meta":{"raw":{"variants":["Social posts betray suicide risk with 82–97% precision","AI reads suicide risk in social text—but ethics lag behind","Social media words detect suicidal thoughts, but bias remains","82–97% precision: social posts predict suicide risk, not yet deployable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000357,"raw_usage":{"total_tokens":1922,"prompt_tokens":919,"completion_tokens":1003,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":930}},"tokens_in":535,"tokens_out":1003,"duration_ms":8349,"temperature":1.0,"reasoning_tokens":930,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:05:52.180879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train one of the reviewed models on posts from one platform and time period, then test it on posts from a different platform and time period; if precision and recall fall well below the quoted 82–97% and 71–94% ranges, the transferable-screening claim collapses.","supporting_citations":[{"cited_title":"Detection of suicidal ideation on Twitter using machine learning & ensemble approaches,","cited_arxiv_id":null,"evidence_quote":"Supplies the high-accuracy Twitter Random Forest result (98.7% precision, 98.2% recall) that anchors the quoted performance range."},{"cited_title":"Detection of suicide ideation in social media forums using deep learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the Reddit LSTM-CNN deep learning result (93.8% accuracy) used to support deep-learning detection claims."},{"cited_title":"Deep neural networks detect suicide risk from textual facebook posts,","cited_arxiv_id":null,"evidence_quote":"Supplies large-scale Facebook data and ANN models whose AUC results support real-user post detection."},{"cited_title":"Leveraging Digital Health and Machine Learning Toward Reducing Suicide - From Panacea to Practical Tool,","cited_arxiv_id":null,"evidence_quote":"Carries the paper's caution about practical and ethical hurdles to deploying machine learning for suicide prevention."}],"review_version":1}