Pith. sign in

REVIEW 1 major objections 8 minor 85 references

Victims of tech-abuse seeking help online get relevant but unsafe guidance — phishing links, toxic forum replies, and advice that can destroy evidence — and no channel is consistently safe.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:05 UTC pith:BI2X6LNW

load-bearing objection A valuable victim-centered dataset and framework, but the Reddit arm of the cross-platform comparison is structurally confounded and needs a fix before the headline claims can be trusted. the 1 major comments →

arxiv 2607.21549 v1 pith:BI2X6LNW submitted 2026-07-23 cs.CY cs.SI

Seeking Help in the Digital Age: A Cross-Platform Analysis of Online Support Systems for Technology-Facilitated Abuse Victims

classification cs.CY cs.SI
keywords technology-facilitated abuseonline help-seekingweb search safetypeer-support forumsconversational AItrauma-informed supportphishing exposurevictim-centered evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Technology-facilitated abuse — stalking, harassment, surveillance, and control carried out through phones, social media, trackers, financial apps, and smart devices — drives many victims to the internet for help when formal support feels inaccessible. This paper asks whether that help is actually safe. Using 2,797 real victim-authored questions drawn from a decade of r/Stalking narratives, it simulates help-seeking across Google Search, Reddit, and five conversational AI systems, grading responses on technical quality (relevance, accuracy, actionability, persuasiveness, understandability) and on safety (malicious-link exposure, toxicity, empathy, bias, risk-aware guidance, and support referrals). The central claim is that no channel consistently provides safe, trauma-informed support: more than 65% of victim queries encounter potentially malicious links in search results, over 20% of Reddit discussions contain toxic responses, and even the best AI systems give 'damaging guidance' — technically plausible advice that can destroy evidence or escalate danger — in about one response in five. The stakes: for a population already hesitant to seek formal help, a search result or chatbot answer can shape a safety-critical decision, so the finding reframes the act of seeking help itself as a source of new risk.

Core claim

The paper's central claim is that online help for victims of technology-facilitated abuse can actively introduce new harm, not merely fail to help. Google Search and general-purpose LLMs beat peer forums on technical quality, yet no channel clears the safety bar. Over 65% of victim queries encountered malicious links in Google results, over 20% of Reddit threads contained toxic comments, and even the strongest AI systems produced 'Damaging Guidance' — plausible advice that would destroy evidence or escalate risk — roughly once in five responses. Surprisingly, survivor-support chatbots underperformed general-purpose LLMs across nearly every dimension — a design problem, not a knowledge gap.

What carries the argument

The load-bearing object is the victim-centered query corpus: 2,797 real help-seeking questions extracted from a decade of r/Stalking narratives through LLM-assisted extraction and validated classifiers, spanning 11 technology-misuse categories. Around it the paper builds a Unified Evaluation Framework that grades every response twice — on five technical qualities (relevance, accuracy, actionability, persuasiveness, understandability) and on platform-specific safety characteristics (malicious-link exposure for web results, toxicity for forum threads, and five trauma-informed dimensions for AI systems: empathy and humanization, voice and choice, bias, risk-informed guidance, and support inform

Load-bearing premise

The paper assumes that questions extracted from r/Stalking — one self-selected, English-language community — represent the help-seeking needs of the broader TFA victim population, and that a single rubric can fairly compare long webpages, comment threads, and single-turn chatbot replies.

What would settle it

A reader could rerun the pipeline on queries from a different victim community — say a domestic-violence forum or a non-English support board — using the same rubric. If those queries yielded consistently safe, trauma-informed, and actionable guidance across all three channels, or if phishing-exposure and toxicity rates fell far below the reported 65% and 20%, the claim that online help-seeking systematically endangers TFA victims would be weakened. More narrowly, an independent re-annotation of the 90-pair webpage and 50-pair chatbot validation sets would test whether the cross-platform ranki

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • No single channel is safe to rely on: search covers the most queries but exposes roughly two-thirds of them to malicious links, forums provide community but almost no actionable or fully accurate guidance, and chatbots give the most actionable answers while still producing risky advice in about one in five responses.
  • Sound-sounding advice can be harmful in context: resetting a device, deleting an account, or blocking an abuser is standard security guidance but 'Damaging Guidance' for TFA victims because it destroys evidence and can escalate danger.
  • The least well-supported queries cut across the highest-consequence categories — financial platforms, people-search sites, image/video manipulation, and surveillance/tracking — so targeted work on these categories would address the weakest areas.
  • Because malicious-link exposure was consistent across all technology categories (63–69% of queries), the paper concludes no subgroup of victims is insulated from the risks of online help-seeking.
  • The authors attribute the survivor-chatbots' underperformance to design and evaluation shortcomings rather than missing domain knowledge, implying that safety-centered design is the binding constraint.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the paper is right, a natural next step is an intervention study: add safe-result filtering, evidence-preservation warnings, and crisis-resource inserts to search and chatbot outputs, then measure whether victims' protective actions improve.
  • The pipeline — survivor narratives to query corpus to unified safety grading — transfers to other help-seeking populations such as youth, immigrant survivors, or non-English speakers, who would likely need an expanded misuse taxonomy and locale-specific resources.
  • The headline rates are a snapshot of a fast-moving ecosystem; repeating the measurement periodically would show whether platforms are getting safer or merely changing the form of the harm.
  • One implication the authors leave implicit: victim-support organizations should treat search results and chatbot replies as part of the victim's risk environment — for example by publishing curated, pre-vetted search links and recommended prompts — rather than as neutral information.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 8 minor

Summary. The paper constructs a victim-centered dataset of technology-facilitated abuse (TFA) help-seeking queries by extracting questions from a decade of r/Stalking posts, classifying 11 technology-misuse categories, and simulating those queries across Google Search, existing Reddit comment threads, three general-purpose LLMs, and two domain-specific chatbots. Responses are evaluated on technical dimensions (relevance, accuracy, actionability, persuasiveness, understandability) and platform-specific social/safety dimensions (social-engineering risk, toxicity, empathy, voice & choice, bias, risk-informed guidance, support information). The headline finding is that Google Search and general-purpose LLMs provide considerably more relevant and actionable guidance than Reddit discussions, yet none of the systems consistently provide safe, trauma-informed support; the paper also reports that over 65% of queries encounter potentially malicious secondary URLs, over 20% of Reddit threads contain toxic comments, and domain-specific chatbots underperform general-purpose LLMs.

Significance. If valid, this is a timely and important contribution to the security, HCI, and victim-support literatures. The paper's strengths include a novel victim-authored query dataset, a multi-dimensional evaluation framework co-developed with social-work experts, explicit human validation of several automated classifiers, and public release of code, prompts, and a sample of processed data. The virus-total and Perspective-API measurements, the emphasis on evidence-destruction risks in technical advice, and the comparison of survivor-support chatbots against general-purpose LLMs are useful and falsifiable. However, the cross-platform comparison currently rests on a structural mismatch in how Reddit responses are paired with queries, and several core manually labeled metrics have low reported inter-rater reliability. These issues do not negate the value of the dataset or framework, but they do undermine the specific cross-platform ranking claims as currently stated.

major comments (1)
  1. [Throughout] The paper uses several thresholds and free parameters (e.g., PQCS τ=0.6, VirusTotal two-engine threshold, Perspective API 0.5, relevance validation sample size, understandability grade cutoff) without a sensitivity analysis. None of these is inherently wrong, but the headline percentages (65.5% malicious URLs, 52% Reddit relevance, 20% toxicity) are all threshold-dependent. A short sensitivity appendix showing how these figures vary with reasonable threshold changes would increase confidence in the conclusions.
minor comments (8)
  1. [Abstract and Section 4] The word 'reponses' appears in the full-text abstract; should be 'responses.' Also, the acknowledgments contain 'NationalbScience Foundation' — a typo for 'National Science Foundation.'
  2. [Figure 2b and Figure 13] Figure 2b's x-axis label contains Unicode/LaTeX artifacts ('Total/uni00A0Questions/uni00A0per/uni00A0Post'), and similar artifacts appear elsewhere. These should be cleaned before camera-ready.
  3. [Section 6.1.2] The heading misspells 'Cross-Platform' as 'Cross-Platfrom.'
  4. [Section 6.3.1] The sentence reporting 'κ=0.4, α=0.6' does not specify which annotation task these values refer to. The same sentence mentions two-coder and three-coder annotation; please attach the reliability statistic to the corresponding task and format.
  5. [Section 7.3.1 and Figure 9] For Empathy & Humanization and Voice & Choice, the paper reports 'the distribution of ratings across coders' rather than a single consensus label. The figure's stacked bars use 'Percentage of Coder Ratings'; this is acceptable, but the caption should state that the unit is coder ratings, not responses, to avoid confusion.
  6. [Section 4] The description of the Reddit response collection says that 2,476 of 2,797 queries had at least one associated comment. Since multiple queries can share the same post, the number of unique posts with comments is not reported. Please state both the number of posts and the number of queries, and explain how queries from the same post are treated in the evaluation.
  7. [Section 5.1] The Flesch–Kincaid Grade Level is used as the Understandability measure. The paper acknowledges this in Appendix F, but the main text should note that Flesch–Kincaid is a proxy and does not capture domain-specific jargon (e.g., 'spyware,' 'two-factor authentication'), which could be exactly what makes content inaccessible to victims.
  8. [Section 8] The Discussion states that domain-specific chatbots underperform general-purpose LLMs 'consistent with Prakash et al. [41]' — a 2026 reference. If this work is not yet published or is under review, please mark it as 'in press' or 'manuscript under review' so readers can judge the citation.

Circularity Check

0 steps flagged

No circularity: the paper is an empirical measurement study whose rubric, classifiers, and thresholds are externally anchored or manually validated; the Reddit query-alignment issue is a validity threat, not a derivation loop.

full rationale

The paper's chain is an empirical measurement pipeline rather than a formal derivation. Victim queries are extracted from r/Stalking posts and then used as inputs to Google Search, Reddit threads, and LLM/chatbot systems; quality is assessed with rubrics whose validation sets are human-annotated (e.g., relevance accuracy 87% for webpages/comments and 90% for LLM responses, Table 3). Technical metrics are adapted from external prior work [54]–[61], and social metrics are informed by external literature and consultation with social workers and victim advocates. No parameter is fitted to a subset of the evaluation outcomes and then renamed as a prediction. Thresholds such as VirusTotal >= 2 engines, Perspective > 0.5, and PQCS tau = 0.6 are stated as choices, not fitted to the headline comparisons. The authors cite some prior work from the same research group (e.g., [63]–[66], [74]–[76], [80]), but these citations support standard methodological conventions (toxicity thresholds, URL scanning practices, related analyses) and are not load-bearing for the central cross-platform claim. The skeptical concern that Reddit responses are the original comment threads rather than replies to the isolated extracted queries is a real external-validity and construct-matching limitation, but it is not circular: relevance is an empirical classification of thread content with respect to each query, and the finding that only 52% of Reddit queries received relevant responses could in principle have gone the other way; nothing in the evaluation definition forces the observed ranking by construction. The 'LLM-as-a-judge' components are validated against human consensus labels before scaling, so they do not reduce to the model's own outputs. No self-definitional step, fitted-input-called-prediction step, load-bearing self-citation, imported uniqueness theorem, ansatz-smuggling via citation, or renaming of a known result was found. Score 0.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical or conceptual entities; it uses an existing taxonomy of technology-misuse categories and a new dataset. The central claims rest on domain assumptions about representativeness, validity of proxy measures (LLM-as-a-judge, VirusTotal, Perspective API), and several hand-chosen thresholds. These are not circular, but they frame the interpretation of the empirical findings.

free parameters (5)
  • VirusTotal malicious-label threshold = ≥2 antivirus flags
    A URL is labeled malicious if at least two VirusTotal engines flag it. This threshold is chosen by hand following prior work and directly affects the 65.5% malicious-link finding (Section 7.1).
  • Perspective API toxicity threshold = >0.5
    Attribute scores above 0.5 are treated as positive for toxicity. This threshold affects the 'over 20% of Reddit discussions contain toxic responses' finding (Section 7.2).
  • PQCS similarity threshold τ = 0.6
    Threshold for retaining post sentences in the post-question context similarity metric, selected by manual calibration (Appendix D). It influences the choice of Llama3.3 over GPT-4 for question extraction, which shapes the downstream query dataset.
  • LLM evaluation sample size = 50 queries
    Only 50 of the 2,797 victim queries were used to evaluate LLMs/chatbots (250 QA pairs), for feasibility. All LLM/chatbot comparisons, including the underperformance of domain-specific chatbots, rest on this small sample (Section 4).
  • Understandability threshold = grades 9-12
    Responses at or below grades 9-12 on the Flesch-Kincaid scale are treated as understandable for general audiences; affects the interpretation of readability results (Appendix F).
axioms (6)
  • domain assumption r/Stalking is a representative source of TFA victim help-seeking narratives.
    The entire query dataset is derived from this single subreddit (Section 3 'Data Collection'), and all downstream findings depend on its representativeness.
  • domain assumption LLM-extracted questions faithfully preserve victims' information needs.
    Question extraction is validated indirectly via PQCS semantic similarity against posts, not against a gold standard of 'true' help-seeking questions (Section 3.2, Appendix D).
  • domain assumption LLM-as-a-judge, validated on small samples, is a valid proxy for large-scale human evaluation.
    Relevance, actionability, and persuasiveness are scored by an ensemble of three LLM judges validated on 90-100 manually annotated examples. The validation samples are small and inter-coder agreement is moderate (Section 6).
  • domain assumption Reference answers for technical accuracy are correct and complete.
    Technical accuracy is assessed against manually constructed reference answers grounded in authoritative sources; if the references are wrong or incomplete, all accuracy labels are affected (Section 6.2.1).
  • domain assumption A unified evaluation rubric can be applied comparably across webpages, comments, and chatbot responses.
    The framework (Section 5) applies the same technical metrics to long-form webpages, fragmented comment threads, and single-turn chatbot replies; differences in format may confound quality comparisons.
  • domain assumption VirusTotal and Perspective API provide valid indicators of maliciousness and toxicity.
    Social-engineering risk and toxicity are operationalized using these external tools (Sections 7.1 and 7.2); their false-positive/false-negative rates are not independently assessed here.

pith-pipeline@v1.3.0-alltime-deepseek · 27079 in / 14271 out tokens · 134544 ms · 2026-08-01T07:05:23.737765+00:00 · methodology

0 comments
read the original abstract

Technology-facilitated abuse (TFA), the use of digital technologies to stalk, harass, monitor or threaten others, has become a pervasive form of interpersonal harm. As victims turn to online sources for guidance, responses can shape how they assess risks, interpret abuse, and choose protective actions. We present a large-scale evaluation of online support for TFA victims across three channels: web search, peer-support forums, and conversational AI systems. Drawing on a decade of victim narratives from r/Stalking, we use qualitative coding and supervised classifiers to construct a dataset of TFA queries spanning 11 categories of technology misuse. We simulate these queries across the three channels and evaluate responses using a unified framework spanning technical, social, and safety dimensions. The framework assesses relevance, accuracy, actionability, persuasiveness, and understandability, alongside platform risks and support characteristics, including social-engineering risk, toxicity, empathy, bias, risky guidance, and support information. We build and validate automated classifiers to scale the evaluation. Our findings reveal differences in support quality across platforms. Google Search and general-purpose LLMs provide more relevant and actionable guidance than Reddit discussions, yet none consistently provide safe, trauma-informed support. More than 65% of victim queries encounter potentially malicious links in search results, over 20% of Reddit discussions contain toxic responses, and conversational AI systems frequently fail to provide risk-aware guidance or concrete support resources. Surprisingly, domain-specific survivor-support chatbots underperform general-purpose LLMs across most dimensions. These findings expose weaknesses in digital support for TFA victims and highlight the need for safety-centered design, evaluation, and deployment of future support technologies.

Figures

Figures reproduced from arXiv: 2607.21549 by Minjaal Raval, Mohit Singhal, Morgan PettyJohn, Nowshin Tabassum, Rachel Voth Schrag, Shirin Nilizadeh, Solomon G. Dandekar, Tim Ryan.

Figure 1
Figure 1. Figure 1: Overview of our Methodology Pipeline safety dimensions, enabling systematic comparison across platforms and categories of technology-facilitated abuse. Data Collection. We focus on the r/Stalking subred￾dit [43] because it is one of the few large, public com￾munities where technology misuse is embedded within real-world interpersonal abuse narratives and help-seeking discussions. We collected Reddit posts … view at source ↗
Figure 2
Figure 2. Figure 2: Distribution and nature of help-seeking questions [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Cross-Platform Relevance QA pairs were sampled uniformly at random without strati￾fication by technology category. Two coders with computer science backgrounds independently annotated all pairs using the relevance rubric. Relevance was defined as whether a response contained technical information or advice that fully or partially addressed the victim’s question; factual and technical correctness were evalu… view at source ↗
Figure 4
Figure 4. Figure 4: Proportion of queries with at least one relevant response grouped by technology type [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Accuracy Evaluation when they provide vague advice or fail to identify concrete technical steps victims can take. 6.3.1. Evaluation Methodology. We evaluated Technical Actionability using the same LLM-as-a-Judge ensemble em￾ployed for Relevance (Section 6.1), comprising Llama 3.3, GPT-OSS 20B, and Gemma3:12b with majority voting. Each judge received the victim query, response, and Tech￾nical Actionability … view at source ↗
Figure 6
Figure 6. Figure 6: Actionability evaluation victim’s situation. We use two labels: Persuasive and Not Persuasive. A response is labeled Persuasive if it provides reasoning, justification, reassurance, or empowering lan￾guage supporting the recommended action, e.g “Enable two￾factor authentication; it prevents your ex from logging in even if they know your password” is Persuasive because it links the recommendation to a clear… view at source ↗
Figure 7
Figure 7. Figure 7: Persuasiveness Evaluation phishing malicious suspicious malware spam not recommended 0 20 40 60 80 100 Malicious URLs (%) 64.4 47.6 21.1 18.8 5.2 2.3 (a) Malicious URL Categories PROFANITY THREAT TOXICITY SEVERE_TOXICITY INSULT IDENTITY_ATTACK 0 20 40 60 80 100 Queries (%) 21.0 3.3 25.9 0.8 18.9 0.5 Has Toxicity No Toxicity (b) Type of toxicity [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Social Engineering in Google search websites (a) [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Social metrics evaluation of LLM responses. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Example questions extracted by Llama3.3 and [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Examples of victim, abuser, TFA, and non-TFA queries extracted from Reddit posts. [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Co-occurrence of technology-misuse categories [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 14
Figure 14. Figure 14: Accuracy of LLMs across tech-misuse categories. [PITH_FULL_IMAGE:figures/full_fig_p019_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: CDF of Malicious Urls and Toxic Comments in [PITH_FULL_IMAGE:figures/full_fig_p019_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

85 extracted references · 3 linked inside Pith

  1. [1]

    “a stalker’s paradise

    D. Freed, J. Palmer, D. Minchala, K. Levy, T. Ristenpart, and N. Dell, ““a stalker’s paradise” how intimate partner abusers exploit technol- ogy,” inProceedings of the 2018 CHI conference on human factors in computing systems, 2018, pp. 1–13

  2. [2]

    Intimate partner violence, technology, and stalking,

    C. Southworth, J. Finn, S. Dawson, C. Fraser, and S. Tucker, “Intimate partner violence, technology, and stalking,”Violence against women, vol. 13, no. 8, pp. 842–856, 2007

  3. [3]

    Technology-facilitated violence and abuse: International perspectives and experiences,

    J. Bailey, N. Henry, and A. Flynn, “Technology-facilitated violence and abuse: International perspectives and experiences,” inThe emer- ald international handbook of technology-facilitated violence and abuse. Emerald Publishing Limited, 2021, pp. 1–17

  4. [4]

    ‘i feel like we’re really behind the game’: perspectives of the united kingdom’s intimate partner violence support sector on the rise of technology-facilitated abuse,

    L. M. Tanczer, I. L ´opez-Neira, and S. Parkin, “‘i feel like we’re really behind the game’: perspectives of the united kingdom’s intimate partner violence support sector on the rise of technology-facilitated abuse,”Journal of gender-based violence, vol. 5, no. 3, pp. 431–450, 2021

  5. [5]

    Digital technologies and intimate partner violence: A qual- itative analysis with multiple stakeholders,

    D. Freed, J. Palmer, D. E. Minchala, K. Levy, T. Ristenpart, and N. Dell, “Digital technologies and intimate partner violence: A qual- itative analysis with multiple stakeholders,”Proceedings of the ACM on human-computer interaction, vol. 1, no. CSCW, pp. 1–22, 2017

  6. [6]

    ” i really just leaned on my community for support

    N. Gupta, K. Walsh, S. Das, and R. Chatterjee, “” i really just leaned on my community for support”: Barriers, challenges, and coping mechanisms used by survivors of{Technology-Facilitated}abuse to seek social support,” in33rd USENIX Security Symposium (USENIX Security 24), 2024, pp. 4981–4998

  7. [7]

    Help-seeking and coping strategies for technology-facilitated abuse experienced by youth,

    D. Freed, S. Consolvo, D. Cosley, P. G. Kelley, E. Ricart, K. Thomas, and N. N. Bazarova, “Help-seeking and coping strategies for technology-facilitated abuse experienced by youth,”Proceedings of the ACM on Human-Computer Interaction, vol. 9, no. 2, pp. 1–25, 2025

  8. [8]

    Tech abuse personas: Exploring help-seeking behaviours and support needs of victim/survivors of technology-facilitated abuse,

    M. Janickyj and L. M. Tanczer, “Tech abuse personas: Exploring help-seeking behaviours and support needs of victim/survivors of technology-facilitated abuse,” inProceedings of the Extended Ab- stracts of the CHI Conference on Human Factors in Computing Systems, 2025, pp. 1–11

  9. [9]

    Disclosure decisions and help-seeking experiences amongst victim-survivors of non-consensual intimate image distribution,

    G. Mclocklin, B. Kellezi, C. Stevenson, and J. Mackay, “Disclosure decisions and help-seeking experiences amongst victim-survivors of non-consensual intimate image distribution,”Victims & Offenders, vol. 20, no. 7, pp. 1258–1284, 2025

  10. [10]

    The web of abuse: A comprehensive analysis of online resource in the context of technology-enabled intimate partner surveillance,

    M. Almansoori, M. Islam, S. Ghosh, M. Mondal, and R. Chatterjee, “The web of abuse: A comprehensive analysis of online resource in the context of technology-enabled intimate partner surveillance,” in2024 IEEE 9th European Symposium on Security and Privacy (EuroS&P). IEEE, 2024, pp. 773–789

  11. [11]

    Designing and evaluating a chatbot for survivors of image-based sexual abuse,

    W. Maeng and J. Lee, “Designing and evaluating a chatbot for survivors of image-based sexual abuse,” inProceedings of the 2022 CHI conference on human factors in computing systems, 2022, pp. 1–21

  12. [12]

    Empathy, bias, and data responsibility: Evaluating ai chatbots for gender-based violence support,

    B. Sanz, M. Lopez-Belloso, and A. Izaguirre Choperena, “Empathy, bias, and data responsibility: Evaluating ai chatbots for gender-based violence support,”Frontiers in Political Science, vol. 7, p. 1631881, 2025

  13. [13]

    {Anti-Privacy}and {Anti-Security}advice on{TikTok}: Case studies of{Technology- Enabled}surveillance and control in intimate partner and{Parent- Child}relationships,

    M. Wei, E. Zeng, T. Kohno, and F. Roesner, “{Anti-Privacy}and {Anti-Security}advice on{TikTok}: Case studies of{Technology- Enabled}surveillance and control in intimate partner and{Parent- Child}relationships,” inEighteenth Symposium on Usable Privacy and Security (SOUPS 2022), 2022, pp. 447–462

  14. [14]

    Online conversations about abuse: Responses to ipv survivors from support communities,

    J. B. Whiting, B. N. Davies, B. C. Eisert, A. B. Witting, and S. R. Anderson, “Online conversations about abuse: Responses to ipv survivors from support communities,”Journal of family violence, vol. 38, no. 5, pp. 791–801, 2023

  15. [15]

    Defining and conceptualizing technology-facilitated abuse (“tech abuse

    N. Koukopoulos, M. Janickyj, and L. M. Tanczer, “Defining and conceptualizing technology-facilitated abuse (“tech abuse”): Findings of a global delphi study,”Journal of Interpersonal Violence, p. 08862605241310465, 2025

  16. [16]

    The spyware used in intimate partner violence,

    R. Chatterjee, P. Doerfler, H. Orgad, S. Havron, J. Palmer, D. Freed, K. Levy, N. Dell, D. McCoy, and T. Ristenpart, “The spyware used in intimate partner violence,” in2018 IEEE Symposium on Security and Privacy (SP). IEEE, 2018, pp. 441–458

  17. [17]

    Clinical computer security for victims of intimate partner violence,

    S. Havron, D. Freed, R. Chatterjee, D. McCoy, N. Dell, and T. Ris- tenpart, “Clinical computer security for victims of intimate partner violence,” in28th USENIX Security Symposium (USENIX Security 19), 2019, pp. 105–122

  18. [18]

    Domestic violence and information communication technologies,

    J. P. Dimond, C. Fiesler, and A. S. Bruckman, “Domestic violence and information communication technologies,”Interacting with com- puters, vol. 23, no. 5, pp. 413–421, 2011

  19. [19]

    Content moderation on social media: Social and com- putational standards and implications,

    M. Singhal, “Content moderation on social media: Social and com- putational standards and implications,” 2024

  20. [20]

    Deepfake technology and gender-based violence: A scoping review,

    L. Lazard, R. Capdevila, E. L. Turley, K. Gilfoyle, and N. Stavropoulou, “Deepfake technology and gender-based violence: A scoping review,”Trauma, Violence, & Abuse, p. 15248380251384271, 2025

  21. [21]

    Characterizing the {MrDeepFakes}sexual deepfake marketplace,

    C. Han, A. Li, D. Kumar, and Z. Durumeric, “Characterizing the {MrDeepFakes}sexual deepfake marketplace,” in34th USENIX Se- curity Symposium (USENIX Security 25), 2025, pp. 5169–5188

  22. [22]

    Deepfakes and digitally altered imagery abuse: A cross-country exploration of an emerging form of image-based sexual abuse,

    A. Flynn, A. Powell, A. J. Scott, and E. Cama, “Deepfakes and digitally altered imagery abuse: A cross-country exploration of an emerging form of image-based sexual abuse,”The British Journal of Criminology, vol. 62, no. 6, pp. 1341–1358, 2022

  23. [23]

    Technology-facilitated sexual violence: Reflections on the concept,

    A. Powell, “Technology-facilitated sexual violence: Reflections on the concept,” inRape. Routledge, 2022, pp. 143–158

  24. [24]

    Abuse vectors: A framework for conceptualizing{IoT- Enabled}interpersonal abuse,

    S. Stephenson, M. Almansoori, P. Emami-Naeini, D. Y . Huang, and R. Chatterjee, “Abuse vectors: A framework for conceptualizing{IoT- Enabled}interpersonal abuse,” in32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 69–86

  25. [25]

    ” it’s the equivalent of feeling like you’re in{Jail

    S. Stephenson, M. Almansoori, P. Emami-Naeini, and R. Chatterjee, “” it’s the equivalent of feeling like you’re in{Jail”}: Lessons from firsthand and secondhand accounts of{IoT-Enabled}intimate partner abuse,” in32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 105–122

  26. [26]

    A global survey of android dual-use applications used in intimate partner surveillance,

    M. Almansoori, A. Gallardo, J. Poveda, A. Ahmed, and R. Chatterjee, “A global survey of android dual-use applications used in intimate partner surveillance,”Proceedings on Privacy Enhancing Technolo- gies, 2022

  27. [27]

    The{Digital-Safety}risks of financial technologies for survivors of intimate partner violence,

    R. Bellini, K. Lee, M. A. Brown, J. Shaffer, R. Bhalerao, and T. Ristenpart, “The{Digital-Safety}risks of financial technologies for survivors of intimate partner violence,” in32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 87–104

  28. [28]

    Abusability of automation apps in intimate partner vi- olence,

    S. Zhang, P. Chung, J. Vervelde, N. Korapati, R. Chatterjee, and K. Fawaz, “Abusability of automation apps in intimate partner vi- olence,” in34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 41–60

  29. [29]

    Gender-based violence and technology- enabled coercive control in seattle: Challenges & opportunities,

    D. Cuomo and N. Dolci, “Gender-based violence and technology- enabled coercive control in seattle: Challenges & opportunities,” TECC Whitepaper Series, 2019

  30. [30]

    A digital safety dilemma: Analysis of computer-mediated computer security interventions for intimate partner violence during covid-19,

    E. Tseng, D. Freed, K. Engel, T. Ristenpart, and N. Dell, “A digital safety dilemma: Analysis of computer-mediated computer security interventions for intimate partner violence during covid-19,” inPro- ceedings of the 2021 CHI Conference on Human Factors in Comput- ing Systems, 2021, pp. 1–17

  31. [31]

    Mavs end tech abuse clinic,

    T. U. of Texas at Arlington, “Mavs end tech abuse clinic,” https: //www.mavsetalab.uta.edu/

  32. [32]

    Madison tech clinic,

    U. of Wisconsin-Madison, “Madison tech clinic,” https://techclinic. cs.wisc.edu/

  33. [33]

    Technology abuse clinics for survivors of intimate partner violence

    L. Ramjit, “Technology abuse clinics for survivors of intimate partner violence.” Santa Clara, CA: USENIX Association, Jan. 2023

  34. [34]

    Is it a crime? cyberstalking victims’ reasons for not reporting to law enforcement,

    E. R. Fissel, “Is it a crime? cyberstalking victims’ reasons for not reporting to law enforcement,”Social Sciences, vol. 12, no. 12, p. 659, 2023

  35. [35]

    Help-seeking from websites and police in the aftermath of technology-facilitated victim- ization,

    D. A. Colburn, D. Finkelhor, and H. A. Turner, “Help-seeking from websites and police in the aftermath of technology-facilitated victim- ization,”Journal of interpersonal violence, vol. 38, no. 21-22, pp. 11 642–11 665, 2023

  36. [36]

    Examining the supports and advice that women with intimate partner violence experience received in online health communities: text mining approach,

    V . Hui, M. Eby, R. E. Constantino, H. Lee, J. Zelazny, J. C. Chang, D. He, and Y . J. Lee, “Examining the supports and advice that women with intimate partner violence experience received in online health communities: text mining approach,”Journal of medical internet research, vol. 25, p. e48607, 2023

  37. [37]

    Ai-ruth,

    N. D. V . Hotline, “Ai-ruth,” https://ruth.thehotline.org/

  38. [38]

    Hopechat ai,

    D. Shelter Org, “Hopechat ai,” https://www.domesticshelters.org/ hope-chat-ai

  39. [39]

    Aimee ai,

    A. Wintemute and S. Nichols, “Aimee ai,” https://www.aimeesays. com/en/home

  40. [40]

    Empathy, bias, and data responsibility: evaluating ai chatbots for gender-based violence support,

    B. Sanz Urquijo, M. L ´opez Belloso, and A. Izaguirre-Choperena, “Empathy, bias, and data responsibility: evaluating ai chatbots for gender-based violence support,”Frontiers in Political Science, vol. 7, p. 1631881, 2025

  41. [41]

    Assessing llm response quality in the context of technology- facilitated abuse,

    V . Prakash, M. Almansoori, D. Hu, R. Chatterjee, and D. Y . Huang, “Assessing llm response quality in the context of technology- facilitated abuse,” 2026

  42. [42]

    Ai-facilitated coercive control: An experimental study,

    H. Kim, T. Ristenpart, and N. Dell, “Ai-facilitated coercive control: An experimental study,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, pp. 1–16

  43. [43]

    Stalking subreddit,

    “Stalking subreddit,” https://www.reddit.com/r/Stalking/

  44. [44]

    Python reddit api wrapper devel- opment,

    P. R. A. W. Development, “Python reddit api wrapper devel- opment,” https://asyncpraw.readthedocs.io/en/stable/code overview/ models/submission.html

  45. [45]

    Reddit comments/submissions 2005-06 to 2025-06

    R. stuck in the matrix, Watchful1, “Reddit comments/submissions 2005-06 to 2025-06.” [Online]. Available: https://academictorrents. com/details/30dee5f0406da7a353aff6a8caa2d54fd01f2ca1

  46. [46]

    B. G. Glaser and A. L. Strauss,Discovery of grounded theory: Strategies for qualitative research. Routledge, 2017

  47. [47]

    A note on the interpretation of weighted kappa and its relations to other rater agreement statistics for metric scales,

    C. Schuster, “A note on the interpretation of weighted kappa and its relations to other rater agreement statistics for metric scales,” Educational and Psychological Measurement, vol. 64, no. 2, pp. 243– 253, 2004

  48. [48]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,”Advances in neural information pro- cessing systems, vol. 33, pp. 1877–1901, 2020

  49. [49]

    Hugging Face, https://huggingface.co/sentence-transformers/ all-MiniLM-L6-v2

  50. [50]

    Sentence-bert: Sentence embeddings using siamese bert-networks,

    N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” inProceedings of the 2019 confer- ence on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), 2019, pp. 3982–3992

  51. [51]

    googlesearch-python (version 1.3.0),

    N. Vikramaditya, “googlesearch-python (version 1.3.0),” PyPI, 2025, https://pypi.org/project/googlesearch-python/ [Accessed: 2025-08-20]

  52. [52]

    Selenium,

    Baiju Muthukadan, “Selenium,” https://selenium-python.readthedocs. io/

  53. [53]

    Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,

    A. Barbaresi, “Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,” in Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations. Association for Computational Ling...

  54. [54]

    Ragas: Automated evaluation of retrieval augmented generation,

    S. Es, J. James, L. E. Anke, and S. Schockaert, “Ragas: Automated evaluation of retrieval augmented generation,” inProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 2024, pp. 150– 158

  55. [55]

    MEMERAG: A multilingual end-to-end meta- evaluation benchmark for retrieval augmented generation,

    M. A. Cruz Bland ´on, J. Talur, B. Charron, D. Liu, S. Mansour, and M. Federico, “MEMERAG: A multilingual end-to-end meta- evaluation benchmark for retrieval augmented generation,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vie...

  56. [56]

    A comprehensive quality evaluation of security and privacy advice on the web,

    E. M. Redmiles, N. Warford, A. Jayanti, A. Koneru, S. Kross, M. Morales, R. Stevens, and M. L. Mazurek, “A comprehensive quality evaluation of security and privacy advice on the web,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 89–108

  57. [57]

    Generating effective answers to people’s everyday cyber- security questions: An initial study,

    A. Balaji, L. Duesterwald, I. Yang, A. Priyanshu, C. Alfieri, and N. Sadeh, “Generating effective answers to people’s everyday cyber- security questions: An initial study,” inInternational Conference on Web Information Systems Engineering. Springer, 2024, pp. 363–379

  58. [58]

    Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems,

    Y . S. Low, M. L. Jackson, R. J. Hyde, R. E. Brown, N. M. Sanghavi, J. D. Baldwin, C. W. Pike, J. Muralidharan, G. Hui, N. Alexander et al., “Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems,”Digital Health, vol. 11, p. 20552076251348850, 2025

  59. [59]

    Measuring what matters: Developing human-centered legal q-and-a quality standards through multi-stakeholder research,

    M. Hagan, “Measuring what matters: Developing human-centered legal q-and-a quality standards through multi-stakeholder research,” Available at SSRN 5146722, 2024

  60. [60]

    Measuring and improving persuasiveness of large language models,

    S. Singh, Y . Singla, H. Si, and B. Krishnamurthy, “Measuring and improving persuasiveness of large language models,” inInterna- tional Conference on Learning Representations, vol. 2025, 2025, pp. 90 267–90 322

  61. [61]

    Comparative evaluation of large language models in explaining radiology reports: Expert assessment of read- ability, understandability, and communication features,

    A. Bozer and Y . Pekc ¸evik, “Comparative evaluation of large language models in explaining radiology reports: Expert assessment of read- ability, understandability, and communication features,”Insights into Imaging, vol. 16, no. 1, pp. 1–10, 2025

  62. [62]

    Flesch reading ease and the flesch kin- caid grade level,

    D. Child, “Flesch reading ease and the flesch kin- caid grade level,” https://readable.com/readability/ flesch-reading-ease-flesch-kincaid-grade-level/

  63. [63]

    Users’ behavioral and emotional response to toxicity in twitter conversations,

    A. Aleksandric, S. S. Roy, H. Pankaj, G. M. Wilson, and S. Nilizadeh, “Users’ behavioral and emotional response to toxicity in twitter conversations,” inProceedings of the International AAAI Conference on Web and Social Media, vol. 18, 2024, pp. 29–42

  64. [64]

    Are these comments triggering? predicting triggers of toxicity in online discus- sions,

    H. Almerekhi, H. Kwak, J. Salminen, and B. J. Jansen, “Are these comments triggering? predicting triggers of toxicity in online discus- sions,” inProceedings of the web conference 2020, 2020, pp. 3033– 3040

  65. [65]

    Under- standing the bystander effect on toxic twitter conversations,

    A. Aleksandric, M. Singhal, A. Groggel, and S. Nilizadeh, “Under- standing the bystander effect on toxic twitter conversations,”arXiv preprint arXiv:2211.10764, 2022

  66. [66]

    User engagement and the toxicity of tweets,

    N. Salehabadi, A. Groggel, M. Singhal, S. S. Roy, and S. Nilizadeh, “User engagement and the toxicity of tweets,”arXiv preprint arXiv:2211.03856, 2022

  67. [67]

    6 key principles of a trauma-informed approach,

    R. Frydman, “6 key principles of a trauma-informed approach,”Cent Integr Heal Solut Trauma, 2020

  68. [68]

    Trauma-informed com- puting: Towards safer technology experiences for all,

    J. X. Chen, A. McDonald, Y . Zou, E. Tseng, K. A. Roundy, A. Tamer- soy, F. Schaub, T. Ristenpart, and N. Dell, “Trauma-informed com- puting: Towards safer technology experiences for all,” inProceedings of the 2022 CHI conference on human factors in computing systems, 2022, pp. 1–20

  69. [69]

    Designing chatbots to support victims and survivors of domestic abuse,

    R. B. Saglam, J. R. Nurse, and L. Sugiura, “Designing chatbots to support victims and survivors of domestic abuse,”arXiv preprint arXiv:2402.17393, 2024

  70. [70]

    Clinic to end tech abuse(ceta),

    Cornell Tech, “Clinic to end tech abuse(ceta),” https://ceta.tech. cornell.edu/resources, 2018

  71. [71]

    Safety net project,

    N. N. to End Domestic Violence, “Safety net project,” https://www. techsafety.org/

  72. [72]

    No more directory,

    T. U. Nations and T. W. Bank, “No more directory,” https:// nomoredirectory.org/usa/

  73. [73]

    Total, https://docs.virustotal.com/reference/overview

    V . Total, https://docs.virustotal.com/reference/overview

  74. [74]

    {DarkGram}: A{Large-Scale}analysis of cybercriminal activity channels on telegram,

    S. S. Roy, E. P. Vafa, K. Khanmohamaddi, and S. Nilizadeh, “{DarkGram}: A{Large-Scale}analysis of cybercriminal activity channels on telegram,” in34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 4839–4858

  75. [75]

    Learning from censored experiences: Social media discussions around censorship circumvention technologies,

    E. P. Vafa, M. Singhal, P. Thota, and S. S. Roy, “Learning from censored experiences: Social media discussions around censorship circumvention technologies,” in2025 IEEE Symposium on Security and Privacy (SP). IEEE, 2025, pp. 1325–1343

  76. [76]

    Cy- bersecurity misinformation detection on social media: Case studies on phishing reports and zoom’s threat,

    M. Singhal, N. Kumarswamy, S. Kinhekar, and S. Nilizadeh, “Cy- bersecurity misinformation detection on social media: Case studies on phishing reports and zoom’s threat,” inProceedings of the Inter- national AAAI Conference on Web and Social Media, vol. 17, 2023, pp. 796–807

  77. [77]

    Perspective api,

    Jigsaw Google, “Perspective api,” https://www.perspectiveapi.com/, 2021

  78. [78]

    Exploring the magnitude and effects of media influence on reddit moderation,

    H. Habib and R. Nithyanand, “Exploring the magnitude and effects of media influence on reddit moderation,” inProceedings of the international AAAI conference on web and social media, vol. 16, 2022, pp. 275–286

  79. [79]

    Re- altoxicityprompts: Evaluating neural toxic degeneration in language models,

    S. Gehman, S. Gururangan, M. Sap, Y . Choi, and N. A. Smith, “Re- altoxicityprompts: Evaluating neural toxic degeneration in language models,” inFindings of the association for computational linguistics: EMNLP 2020, 2020, pp. 3356–3369

  80. [80]

    Causal insights into parler’s content moderation shift: Effects on toxicity and factuality,

    N. Kumarswamy, M. Singhal, and S. Nilizadeh, “Causal insights into parler’s content moderation shift: Effects on toxicity and factuality,” inProceedings of the ACM on Web Conference 2025, 2025, pp. 3762– 3771

Showing first 80 references.