{"id":"69f6106a-54e7-4f35-8f36-a239a3ae79c8","arxiv_id":"2607.21549","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 2,797 victim-authored queries, Google and general-purpose LLMs outperformed Reddit and domain-specific chatbots on relevance and actionability, but none consistently provided safe, trauma-informed guidance.","lead":"This paper measures how well Google Search, Reddit, and AI chatbots answer questions from technology-abuse victims. It finds that search and general-purpose chatbots give more useful advice than Reddit or specialized survivor chatbots, but none is consistently safe.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reddit comparison is confounded: Reddit 'responses' are the original comment threads, not responses to the extracted queries; multi-question posts are scored against every extracted query, so Reddit is penalized for not answering questions its commenters never saw.","rationale":"The reader's weakest_assumption focused on the representativeness of r/Stalking and the fairness of a unified rubric across structurally different formats. My concern is a specific, concrete instance of that cross-format fairness issue: Reddit comments are not generated in response to the extracted queries, while Google and LLM responses are. This asymmetry can bias the headline cross-platform ranking, which is a central claim. The reader did not explicitly identify this response-source mismatch, so my agreement is only partial. The concern does not overturn the paper's broader conclusions (e.g., that no platform is consistently safe), but it does mean the comparative claim about Reddit underperforming requires a query-matched Reddit evaluation before it can be accepted. Since the reader already recommended CONDITIONAL acceptance, my verdict recommendation remains UNCHANGED, but the condition should now specifically include addressing the Reddit response-query alignment issue.","tokens_in":27492,"tokens_out":6795,"duration_ms":65261,"concrete_test":"Re-run the Reddit relevance and actionability analyses stratified by the number of TFA queries extracted per post. For posts with exactly one extracted TFA query (single-query posts), the comment thread is much more likely to be a response to that query. If the single-query subset shows substantially higher relevance/actionability than the multi-query subset, then the low Reddit performance is partly an artifact of evaluating each extracted query against a thread that was not generated to answer it. Additionally, manually annotate a random sample of 100 comment–query pairs from multi-query posts to determine whether each comment actually addresses the specific extracted query; if a large fraction of comments are post-relevant but not query-relevant, the mismatch is confirmed. If the single-query analysis changes the cross-platform ranking, the headline comparison should be re-run with qu","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is a source mismatch in the Reddit arm of the comparison. Section 4 states that Reddit responses are 'the comment threads associated with each r/Stalking post,' while Google and LLM systems receive the exact extracted victim query. Section 3.2/Figure 2b show that a single post can yield multiple extracted questions (up to 7). In Section 6.1.1, the relevance pipeline evaluates each extracted query against the same full comment thread. Thus, for a post with N queries, every query inherits the same thread that was written in response to the post as a whole, not to any isolated query. A comment that answers one question will be scored as non-relevant for the other N-1 questions, and no commenter ever saw those isolated queries. This structurally disadvantages Reddit in the cross-platform comparison and can explain part of the headline finding that only 52% of Reddit queries receive relevant responses (Figure 3a) and that Reddit is the least actionable (1% actionable, Figure 6d). The transfer of the relevance classifier from webpages to Reddit comments without a dedicated validation set (Section 6.1.1) adds further noise, but the primary issue is that the Reddit responses were not generated by the queries being evaluated. Without query-aligned Reddit responses, the claim that 'Google Search and general-purpose LLMs provide considerably more relevant and actionable guidance than Reddit discussions' (Abstract) is not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a victim-centered dataset of technology-facilitated abuse (TFA) help-seeking queries by extracting questions from a decade of r/Stalking posts, classifying 11 technology-misuse categories, and simulating those queries across Google Search, existing Reddit comment threads, three general-purpose LLMs, and two domain-specific chatbots. Responses are evaluated on technical dimensions (relevance, accuracy, actionability, persuasiveness, understandability) and platform-specific social/safety dimensions (social-engineering risk, toxicity, empathy, voice & choice, bias, risk-informed guidance, support information). The headline finding is that Google Search and general-purpose LLMs provide considerably more relevant and actionable guidance than Reddit discussions, yet none of the systems consistently provide safe, trauma-informed support; the paper also reports that over 65% of queries encounter potentially malicious secondary URLs, over 20% of Reddit threads contain toxic comments, and domain-specific chatbots underperform general-purpose LLMs.","tokens_in":27779,"tokens_out":6758,"duration_ms":74413,"significance":"If valid, this is a timely and important contribution to the security, HCI, and victim-support literatures. The paper's strengths include a novel victim-authored query dataset, a multi-dimensional evaluation framework co-developed with social-work experts, explicit human validation of several automated classifiers, and public release of code, prompts, and a sample of processed data. The virus-total and Perspective-API measurements, the emphasis on evidence-destruction risks in technical advice, and the comparison of survivor-support chatbots against general-purpose LLMs are useful and falsifiable. However, the cross-platform comparison currently rests on a structural mismatch in how Reddit responses are paired with queries, and several core manually labeled metrics have low reported inter-rater reliability. These issues do not negate the value of the dataset or framework, but they do undermine the specific cross-platform ranking claims as currently stated.","major_comments":[{"comment":"The paper uses several thresholds and free parameters (e.g., PQCS τ=0.6, VirusTotal two-engine threshold, Perspective API 0.5, relevance validation sample size, understandability grade cutoff) without a sensitivity analysis. None of these is inherently wrong, but the headline percentages (65.5% malicious URLs, 52% Reddit relevance, 20% toxicity) are all threshold-dependent. A short sensitivity appendix showing how these figures vary with reasonable threshold changes would increase confidence in the conclusions.","section":"Throughout"}],"minor_comments":[{"comment":"The word 'reponses' appears in the full-text abstract; should be 'responses.' Also, the acknowledgments contain 'NationalbScience Foundation' — a typo for 'National Science Foundation.'","section":"Abstract and Section 4"},{"comment":"Figure 2b's x-axis label contains Unicode/LaTeX artifacts ('Total/uni00A0Questions/uni00A0per/uni00A0Post'), and similar artifacts appear elsewhere. These should be cleaned before camera-ready.","section":"Figure 2b and Figure 13"},{"comment":"The heading misspells 'Cross-Platform' as 'Cross-Platfrom.'","section":"Section 6.1.2"},{"comment":"The sentence reporting 'κ=0.4, α=0.6' does not specify which annotation task these values refer to. The same sentence mentions two-coder and three-coder annotation; please attach the reliability statistic to the corresponding task and format.","section":"Section 6.3.1"},{"comment":"For Empathy & Humanization and Voice & Choice, the paper reports 'the distribution of ratings across coders' rather than a single consensus label. The figure's stacked bars use 'Percentage of Coder Ratings'; this is acceptable, but the caption should state that the unit is coder ratings, not responses, to avoid confusion.","section":"Section 7.3.1 and Figure 9"},{"comment":"The description of the Reddit response collection says that 2,476 of 2,797 queries had at least one associated comment. Since multiple queries can share the same post, the number of unique posts with comments is not reported. Please state both the number of posts and the number of queries, and explain how queries from the same post are treated in the evaluation.","section":"Section 4"},{"comment":"The Flesch–Kincaid Grade Level is used as the Understandability measure. The paper acknowledges this in Appendix F, but the main text should note that Flesch–Kincaid is a proxy and does not capture domain-specific jargon (e.g., 'spyware,' 'two-factor authentication'), which could be exactly what makes content inaccessible to victims.","section":"Section 5.1"},{"comment":"The Discussion states that domain-specific chatbots underperform general-purpose LLMs 'consistent with Prakash et al. [41]' — a 2026 reference. If this work is not yet published or is under review, please mark it as 'in press' or 'manuscript under review' so readers can judge the citation.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important, understudied problem and has a strong empirical core. The main barrier to publication is the structural mismatch in the Reddit comparison: the paper cannot claim that Google and LLMs are 'more relevant and actionable' than Reddit when Reddit responses are not generated by the queries being scored. This is fixable with a re-analysis (single-query posts, post-level scoring, or a new Reddit data collection), but it is load-bearing for the abstract. The low inter-rater reliability on several central metrics should also be addressed head-on. I would not reject the paper; the dataset and framework are valuable even if the cross-platform ranking is downgraded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The work is worth reading: it builds a genuinely new dataset of 2,797 victim-authored TFA queries from a decade of r/Stalking posts, develops an 11-category taxonomy, and evaluates Google Search, Reddit, and five conversational AI systems across technical and social dimensions. The unified framework is a real contribution, as is the effort to validate LLM-based classifiers against human annotations. Some findings look solid because they don't depend on the trickiest comparison: phishing-heavy malicious links in search results, toxicity in a substantial share of Reddit threads, LLMs producing \"damaging guidance\" that could harm victims, and domain-specific survivor chatbots underperforming general-purpose models.\n\nThe soft spot is large and sits at the center of the paper's main claim. For Reddit, the \"responses\" are the original comment threads on each post, while Google and the LLMs receive the exact extracted query. Since a single post can yield multiple extracted questions, every question from that post is scored against the same thread, which was written in response to the post as a whole, not to any isolated question. That systematically disadvantages Reddit: a thread that answers one of three extracted questions gets scored as non-relevant for the other two, even though no commenter ever saw them. So the finding that \"Google Search and general-purpose LLMs provide considerably more relevant and actionable guidance than Reddit discussions\" is not yet supported. The fix is conceptually easy: evaluate Reddit at post level, or use the original post text as the query for Reddit, or collect new responses to the extracted queries. But as it stands, the cross-platform comparison is apples-to-oranges.\n\nOther issues are more minor: the LLM/chatbot evaluation uses only 50 queries; inter-rater reliability is moderate for accuracy on webpages and for the social dimensions of bias and risk-informed guidance; the relevance classifier validated on webpages is applied to Reddit comments without a separate validation set; and headline percentages lack confidence intervals. None of these break the study on their own, but they add noise.\n\nWho is this for? Anyone working on digital safety, support infrastructure for survivors, or LLM evaluation in sensitive domains. The dataset and framework deserve to be built on; the Reddit comparison needs repair first. I'd send it to peer review, with a clear request for a revised Reddit methodology.","headline":"A valuable victim-centered dataset and framework, but the Reddit arm of the cross-platform comparison is structurally confounded and needs a fix before the headline claims can be trusted.","tokens_in":28358,"tokens_out":2738,"would_cite":true,"duration_ms":27363,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Victims of tech-abuse seeking help online get relevant but unsafe guidance — phishing links, toxic forum replies, and advice that can destroy evidence — and no channel is consistently safe.","keywords":["technology-facilitated abuse","online help-seeking","web search safety","peer-support forums","conversational AI","trauma-informed support","phishing exposure","victim-centered evaluation"],"falsifier":"A reader could rerun the pipeline on queries from a different victim community — say a domestic-violence forum or a non-English support board — using the same rubric. If those queries yielded consistently safe, trauma-informed, and actionable guidance across all three channels, or if phishing-exposure and toxicity rates fell far below the reported 65% and 20%, the claim that online help-seeking systematically endangers TFA victims would be weakened. More narrowly, an independent re-annotation of the 90-pair webpage and 50-pair chatbot validation sets would test whether the cross-platform ranki","tokens_in":27319,"feed_emoji":"🛡️","tokens_out":22606,"duration_ms":162039,"temperature":0.7,"pith_summary":"Technology-facilitated abuse — stalking, harassment, surveillance, and control carried out through phones, social media, trackers, financial apps, and smart devices — drives many victims to the internet for help when formal support feels inaccessible. This paper asks whether that help is actually safe. Using 2,797 real victim-authored questions drawn from a decade of r/Stalking narratives, it simulates help-seeking across Google Search, Reddit, and five conversational AI systems, grading responses on technical quality (relevance, accuracy, actionability, persuasiveness, understandability) and on safety (malicious-link exposure, toxicity, empathy, bias, risk-aware guidance, and support referrals). The central claim is that no channel consistently provides safe, trauma-informed support: more than 65% of victim queries encounter potentially malicious links in search results, over 20% of Reddit discussions contain toxic responses, and even the best AI systems give 'damaging guidance' — technically plausible advice that can destroy evidence or escalate danger — in about one response in five. The stakes: for a population already hesitant to seek formal help, a search result or chatbot answer can shape a safety-critical decision, so the finding reframes the act of seeking help itself as a source of new risk.","feed_headline":"65% of tech-abuse victims' searches surface malicious links","feed_subtitle":"Study of 2,797 real victim questions finds no channel is consistently safe or trauma-informed.","key_machinery":"The load-bearing object is the victim-centered query corpus: 2,797 real help-seeking questions extracted from a decade of r/Stalking narratives through LLM-assisted extraction and validated classifiers, spanning 11 technology-misuse categories. Around it the paper builds a Unified Evaluation Framework that grades every response twice — on five technical qualities (relevance, accuracy, actionability, persuasiveness, understandability) and on platform-specific safety characteristics (malicious-link exposure for web results, toxicity for forum threads, and five trauma-informed dimensions for AI systems: empathy and humanization, voice and choice, bias, risk-informed guidance, and support inform","core_discovery":"The paper's central claim is that online help for victims of technology-facilitated abuse can actively introduce new harm, not merely fail to help. Google Search and general-purpose LLMs beat peer forums on technical quality, yet no channel clears the safety bar. Over 65% of victim queries encountered malicious links in Google results, over 20% of Reddit threads contained toxic comments, and even the strongest AI systems produced 'Damaging Guidance' — plausible advice that would destroy evidence or escalate risk — roughly once in five responses. Surprisingly, survivor-support chatbots underperformed general-purpose LLMs across nearly every dimension — a design problem, not a knowledge gap.","pith_inferences":["If the paper is right, a natural next step is an intervention study: add safe-result filtering, evidence-preservation warnings, and crisis-resource inserts to search and chatbot outputs, then measure whether victims' protective actions improve.","The pipeline — survivor narratives to query corpus to unified safety grading — transfers to other help-seeking populations such as youth, immigrant survivors, or non-English speakers, who would likely need an expanded misuse taxonomy and locale-specific resources.","The headline rates are a snapshot of a fast-moving ecosystem; repeating the measurement periodically would show whether platforms are getting safer or merely changing the form of the harm.","One implication the authors leave implicit: victim-support organizations should treat search results and chatbot replies as part of the victim's risk environment — for example by publishing curated, pre-vetted search links and recommended prompts — rather than as neutral information."],"forward_implications":["No single channel is safe to rely on: search covers the most queries but exposes roughly two-thirds of them to malicious links, forums provide community but almost no actionable or fully accurate guidance, and chatbots give the most actionable answers while still producing risky advice in about one in five responses.","Sound-sounding advice can be harmful in context: resetting a device, deleting an account, or blocking an abuser is standard security guidance but 'Damaging Guidance' for TFA victims because it destroys evidence and can escalate danger.","The least well-supported queries cut across the highest-consequence categories — financial platforms, people-search sites, image/video manipulation, and surveillance/tracking — so targeted work on these categories would address the weakest areas.","Because malicious-link exposure was consistent across all technology categories (63–69% of queries), the paper concludes no subgroup of victims is insulated from the risks of online help-seeking.","The authors attribute the survivor-chatbots' underperformance to design and evaluation shortcomings rather than missing domain knowledge, implying that safety-centered design is the binding constraint."],"fun_headline_variants":["65% of tech-abuse victims' searches hit malicious links","Survivor chatbots underperform general AI for abuse help","No online abuse support is both safe and trauma-informed","AI gives risky abuse guidance in 1 in 5 victim replies","Google and AI beat Reddit, but none safe for abuse victims"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that questions extracted from r/Stalking — one self-selected, English-language community — represent the help-seeking needs of the broader TFA victim population, and that a single rubric can fairly compare long webpages, comment threads, and single-turn chatbot replies.","fun_headline_variants_meta":{"raw":{"variants":["65% of tech-abuse victims' searches hit malicious links","Survivor chatbots underperform general AI for abuse help","No online abuse support is both safe and trauma-informed","AI gives risky abuse guidance in 1 in 5 victim replies","Google and AI beat Reddit, but none safe for abuse victims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000822,"raw_usage":{"total_tokens":3477,"prompt_tokens":830,"completion_tokens":2647,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2572}},"tokens_in":574,"tokens_out":2647,"duration_ms":20534,"temperature":1.0,"reasoning_tokens":2572,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:05:23.737765+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could rerun the pipeline on queries from a different victim community — say a domestic-violence forum or a non-English support board — using the same rubric. If those queries yielded consistently safe, trauma-informed, and actionable guidance across all three channels, or if phishing-exposure and toxicity rates fell far below the reported 65% and 20%, the claim that online help-seeking systematically endangers TFA victims would be weakened. More narrowly, an independent re-annotation of the 90-pair webpage and 50-pair chatbot validation sets would test whether the cross-platform ranki","supporting_citations":[],"review_version":1}