Pith. sign in

REVIEW 2 major objections 5 minor 74 references

A survey of 49 sources finds fraud-LLM papers omit latency, cost, and calibration evidence.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:03 UTC pith:S5VGKBAI

load-bearing objection Useful, honestly-scoped evidence audit; the headline fraud-vs-moderation imbalance is real in the coded corpus but partly vulnerable to the author's own search-asymmetry caveat. the 2 major comments →

arxiv 2607.13078 v1 pith:S5VGKBAI submitted 2026-07-12 cs.CR cs.AIcs.LG

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

classification cs.CR cs.AIcs.LG
keywords LLM deployment evidencefraud detectiontrust and safetycontent moderationlatency and costcalibrationoperational evaluationsurvey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper is an evidence audit of the public literature on using large language models in fraud detection and adjacent trust-and-safety workflows. It codes 49 operationally relevant sources and finds a sharp split: fraud supplies the most task-specific papers, yet among the 18 fraud and investigation sources none report clean per-decision latency, per-decision dollar cost, or calibration evidence. The moderation subfield, by contrast, contributes most of the explicit deployment evidence, including latency, cost, fairness, and governance signals. The paper offers FORTE, a role-and-evidence frame for locating where an LLM sits in a pipeline, and a five-item minimum deployment-evidence checklist. A sympathetic reader would take away that current public evidence supports selective, hybrid deployments but does not establish that an LLM can replace a real-time fraud engine.

Core claim

On the paper's own terms, the central discovery is an evidence imbalance in the coded corpus: fraud detection is the largest task-specific category, yet none of its 18 sources report clean per-decision latency, per-decision dollar cost, or calibration. Most fraud work reports offline task performance, retrieval gains, or case-study accuracy; only one source has page-verified batch-runtime/deployment evidence, one reports token usage, and three report human or user-facing evidence. The moderation and abuse sources, though fewer, include livestream A/B results, streaming token-budget results, longitudinal audits, and demographic-fairness audits. The paper argues this imbalance means deployment

What carries the argument

The load-bearing organizing device is FORTE, a role-and-evidence frame that locates an LLM in a workflow by five lenses: fraud-first anchoring, operational placement patterns, roles (classifier, retrieval interface, explanation generator, reviewer assistant, agent, feature extractor, escalation component), trust axes (robustness, explanation integrity, selective prediction), and evaluation-gap analysis. The paper's quantitative engine is a manually coded evidence matrix of 49 sources, with strict definitions separating clean per-decision latency from batch runtime, dollar cost from token counts, and calibration from accuracy. The frame lets the survey compare papers across fraud, moderation,

Load-bearing premise

The survey's headline gap rests on the assumption that the coded corpus fairly represents the public literature; the author explicitly notes that a targeted scan for moderation-side fairness audits was not matched by an equally targeted scan for fraud-side latency and cost reporting, so part of the imbalance could reflect search effort rather than reality.

What would settle it

A targeted search of the same venues and time window using fraud-side latency, cost, and calibration terms—mirroring the moderation-side scan—that retrieves even one fraud or investigation paper reporting clean per-decision latency, per-decision dollar cost, or calibration would falsify the claim that none exist.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the evidence imbalance is real, public literature does not currently justify deploying an LLM as a real-time fraud scorer without additional latency, cost, and calibration data.
  • Selective and hybrid architectures—fast conventional scoring with LLMs reserved for hard cases, explanation, escalation, or investigation—are the most defensible near-term option supported by the corpus.
  • Fraud papers that report at least three of latency budget, inference cost, explicit decision threshold, and a human-impact metric would substantially close the gap.
  • The minimum deployment-evidence checklist gives practitioners and auditors a concrete standard for judging operational maturity of LLM-based trust-and-safety systems.
  • Robustness findings from cross-cutting work, especially indirect prompt injection and retrieval poisoning, should be treated as mandatory evaluation layers for fraud and moderation deployments.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The measured fraud–moderation imbalance may partly reflect publication norms: fraud teams in production may withhold latency and cost data for adversarial reasons, so a symmetric search of industry and incident-report channels could surface evidence the academic corpus misses.
  • A natural next step is to treat the five-item checklist as a reporting standard for new fraud-LLM papers; if adopted, the field would generate the Pareto-style cost-benefit evidence the survey says is missing.
  • The FORTE role vocabulary could transfer to adjacent high-stakes domains such as healthcare triage or legal review, where the same question—what role does the model play and what evidence supports it—applies verbatim.
  • Because the paper's counts are frozen at a May 2026 snapshot, its 'no fraud source reports calibration' claim is a moving target; post-window papers already reporting calibration in routing suggest the gap may narrow quickly.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper is a structured narrative survey of 49 operationally relevant sources (18 fraud/investigation, 14 moderation/abuse, 17 cross-cutting robustness) plus 15 contextual references, coding each source for application area, LLM role, adaptation style, deployment pattern, and evidence tags. It introduces FORTE, a role-based organizing frame; identifies four deployment patterns; proposes a five-item minimum deployment-evidence checklist; and lists nine follow-up studies. The headline empirical finding is an evidence imbalance: in the coded corpus, no fraud/investigation source reports clean per-decision latency, per-decision dollar cost, or calibration, while moderation sources contain more explicit operational evidence on latency, cost, governance, and fairness. The paper repeatedly and appropriately limits its claims to the coded corpus and acknowledges major limitations: non-exhaustive search, single-author coding without inter-rater agreement, thin banking-sector coverage, and an asymmetric targeted search in the May 2026 update.

Significance. Because the paper is an evidence audit, its value depends heavily on the representativeness of the corpus and the reliability of the coding. Within those limits, the contribution is useful: the FORTE taxonomy and the minimum deployment-evidence checklist are concrete and practitioner-usable, and the paper avoids the common failure of treating 49 sources as 49 production deployments. The explicit definitions for 'clean latency,' 'dollar cost,' and 'calibration' are a strength, and the claim about the coded corpus is falsifiable if the evidence matrix is made available. The honest limitation section is a credit. However, the central comparative claim (fraud lacks operational evidence, moderation has it) is vulnerable to the acknowledged search asymmetry, and the coding reliability is not independently auditable in the manuscript. These issues are fixable without changing the paper's scope.

major comments (2)
  1. [§3, Limitations (fourth bullet); §§1 and 6] The absence claim and the fraud–moderation imbalance are the paper's central result, but the corpus-construction asymmetry directly threatens that result. The May 2026 update ran a targeted scan for moderation-side fairness audits and production moderation systems, while no equally targeted scan was run for fraud-side latency/cost reporting. Because the paper draws a comparative conclusion (fraud lacks operational evidence, moderation has it), the measured imbalance could be an artifact of search effort rather than a property of the literature. The paper flags this, but the abstract and Section 6 nevertheless state the imbalance as a finding about the public literature ('moderation papers ... include more explicit public evidence'). To make the central claim load-bearing, either (a) run the symmetric targeted search and report the query families and screening rules, or (b) consistently r
  2. [§3, 'Coding scheme' and Table 3] The quantitative heart of the paper is the count of evidence types: 0 clean per-decision latency, 0 dollar cost, 0 calibration among 18 fraud sources. These counts depend on single-author judgments about what counts as 'clean,' 'dollar cost,' and 'calibration.' The paper reports no inter-rater reliability, and the coding definitions, while stated, leave room for judgment (e.g., the boundary between 'batch runtime' and 'per-decision latency'; between 'token-usage evidence' and 'dollar cost'). For a survey whose main empirical claim is an absence, the reader should be able to audit the coding. I recommend making the evidence matrix available as supplementary material (currently it is only mentioned as 'included with the source bundle') and having a second coder code a subset with agreement reported. Without this, the zero counts are a claim about the author's coding, not a machine-checkabl
minor comments (5)
  1. [Abstract and throughout] Typesetting errors with missing spaces appear in the abstract and Section 2 (e.g., 'Weaddressthisquestion', 'mechanisms.Safety', 'TheLLM Security and Privacysurvey'). A careful proofreading pass is needed.
  2. [§5, Pattern counts] The counts (19 sources in four patterns; 13 other) are not mapped to individual references. Since the placement coding drives the pattern analysis, a supplementary table listing each source and its pattern/role would increase transparency and help readers verify the counts.
  3. [§6, Evaluation layers] The text says 'One has page-verified batch-runtime/deployment evidence' without naming the source in the paragraph. Please identify the source so readers can verify the claim without consulting the evidence matrix.
  4. [§7, Signals after the search window] The July 2026 spot check is described, but the query and selection criteria are not provided with the same detail as the May 2026 corpus. Including the same transparency for the spot check would allow readers to assess whether the after-window sources are comparable to the coded corpus.
  5. [References] Several references are incomplete or inconsistent: [15] lists only 'Zeng and Zhu'; [38] and [41] have partial author listings. Please complete all references.

Circularity Check

0 steps flagged

No significant circularity: the central claim is an explicitly scoped, externally coded literature measurement, not a derivative of its own inputs.

full rationale

The paper's central claim is an empirical audit of a coded corpus, not a derivation from assumptions. The 18-fraud/14-moderation counts and the absence of clean per-decision latency, dollar cost, or calibration among fraud sources are coding outcomes; the coding categories are defined upfront (Section 6: "Clean latency means per-decision latency, streaming latency, or an explicit latency budget... Dollar cost means a reported monetary cost... Calibration is counted only when a paper evaluates probability calibration, threshold reliability, or coverage-risk behavior") and the claims are explicitly scoped to the corpus: "Among the 18 fraud and investigation sources, none report clean per-decision latency... None report per-decision dollar cost, and none report calibration." No fitted parameter is renamed as a prediction, no equation is constructed from the target result, and there are no self-citations by the author (no Gabani references appear in the bibliography). The acknowledged limitation—the May 2026 update included "a targeted scan for moderation-side fairness audits and production moderation systems, with no equally targeted scan for fraud-side latency/cost reporting"—is a representativeness threat, not circular reasoning; it weakens the generalization from the coded corpus to the literature but does not make the corpus-internal finding tautological. FORTE is an organizing taxonomy and checklist, not an imported uniqueness theorem or an ansatz smuggled in by citation. Because no load-bearing argument reduces to a self-citation, a definitional equivalence, or a fitted input, the honest finding is no circularity, and the score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

All claims are meta-analytic; no fitted parameters. The ledger lists the methodological assumptions that the evidence-gap finding depends on.

axioms (4)
  • domain assumption The searched corpus (snowball from three neighbor surveys plus targeted queries, 2023–2026) is representative of the LLM fraud/trust-and-safety literature.
    Central claim of evidence absence depends on the corpus being representative; author flags this as a limitation in §3 (search effort may differ across categories).
  • ad hoc to paper The seven-role taxonomy is the right decomposition of LLM placement in workflows.
    The taxonomy is introduced by the author; if another role decomposition were used, the coded counts and gap analysis could change.
  • domain assumption Deployment evidence is primarily latency, per-decision cost, calibration, threshold, explanation integrity, and adversarial pressure.
    The five-item checklist defines what counts as evidence; if these are not the right metrics, the gap finding is less meaningful.
  • domain assumption Single-author coding without inter-rater reliability is accurate enough for the headline counts.
    Author notes no inter-rater agreement statistic; coding consistency was handled by re-coding earlier papers, which risks drift (§3).
invented entities (1)
  • FORTE no independent evidence
    purpose: Organizing frame for the survey: fraud-first anchoring, operational placement patterns, seven LLM roles, trust axes, evaluation-gap analysis.
    A taxonomy/frame, not an empirical entity; no falsifiable handle beyond the paper's own coding.

pith-pipeline@v1.3.0-alltime-deepseek · 16850 in / 9411 out tokens · 80921 ms · 2026-08-02T07:03:02.387214+00:00 · methodology

0 comments
read the original abstract

LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.

Figures

Figures reproduced from arXiv: 2607.13078 by Keyur Gabani.

Figure 1
Figure 1. Figure 1: Survey landscape. Existing surveys cover safety of LLMs, fraud-ML methods, or moderation independently. Applied papers are scattered across venues. This survey organizes the deployment￾evidence gap. tigation, SLM-Mod [7] for specialized smaller models under moderation constraints, and Poi￾sonedRAG [13] or the indirect-prompt-injection competition [35] for current adversarial risk. These are not the whole f… view at source ↗
Figure 2
Figure 2. Figure 2: FORTE organizing frame. Fraud detection is the anchor domain; investigation and compliance are downstream workflow stages; moderation and abuse are comparison areas. Robustness, explanation integrity, and selective prediction are cross-cutting axes. coded as investigation support rather than as one of the three dominant-role retrieval papers. In these studies, the LLM supports evidence synthesis rather tha… view at source ↗
Figure 3
Figure 3. Figure 3: Conceptual deployment stack for a selective architecture. Fast first-pass scoring handles routine decisions; LLMs are reserved for hard cases, explanation, and investigation. Monitoring tracks explanation quality, retrieval integrity, and adversarial behavior. The evaluation-layer coding is intentionally strict. Clean latency means per-decision la￾tency, streaming latency, or an explicit latency budget; ba… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

74 extracted references · 1 canonical work pages

  1. [1]

    Understanding structured financial data with LLMs: A case study on fraud detection, 2025

    Xuwei Tan, Yao Ma, and Xueru Zhang. Understanding structured financial data with LLMs: A case study on fraud detection, 2025. arXiv preprint arXiv:2512.13040

  2. [2]

    FLAG: Fraud detection with LLM-enhanced graph neural network

    Chengdong Yang et al. FLAG: Fraud detection with LLM-enhanced graph neural network. InProceedings of KDD, 2025. doi: 10.1145/3711896.3737220

  3. [3]

    LLM-enhanced self-evolving reinforcement learning for multi-step e-commerce pay- ment fraud risk detection

    BoQu, ZhurongWang, DaisukeYagi, ZhenXu, YangZhao, YinanShan, andFrankZahrad- nik. LLM-enhanced self-evolving reinforcement learning for multi-step e-commerce pay- ment fraud risk detection. InProceedings of ACL (Industry Track), 2025. arXiv preprint arXiv:2509.18719

  4. [4]

    Advanced real-time fraud detection using RAG-based LLMs, 2025

    Gurjot Singh, Prabhjot Singh, and Maninder Singh. Advanced real-time fraud detection using RAG-based LLMs, 2025. arXiv preprint arXiv:2501.15290

  5. [5]

    LLM-assistedauthenticationandfrauddetection,

    EmunahS-S.ChanandAldarC-F.Chan. LLM-assistedauthenticationandfrauddetection,

  6. [6]

    AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts, 2024

    Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien. AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts, 2024. arXiv preprint arXiv:2404.05993. 17

  7. [7]

    SLM-Mod: Small language models surpass LLMs at content moderation

    Xianyang Zhan, Agam Goyal, Yilun Chen, Eshwar Chandrasekharan, and Koustuv Saha. SLM-Mod: Small language models surpass LLMs at content moderation. InProceedings of NAACL, 2025. arXiv preprint arXiv:2410.13155

  8. [8]

    Agentic AI framework for enhancing scam intelligence in digital payments, 2025

    Nitish Jaipuria, Lorenzo Gatto, Zijun Kan, Shankey Poddar, Bill Cheung, Diksha Bansal, Ramanan Balakrishnan, Aviral Suri, and Jose Estevez. Agentic AI framework for enhancing scam intelligence in digital payments, 2025. arXiv preprint arXiv:2508.19932

  9. [9]

    EXPLICATE: Enhancing phishing detection through explainable AI and LLM-powered interpretability, 2025

    Bryan Lim, Roman Huerta, Alejandro Sotelo, Anthonie Quintela, and Priyanka Kumar. EXPLICATE: Enhancing phishing detection through explainable AI and LLM-powered interpretability, 2025. arXiv preprint arXiv:2503.20796

  10. [10]

    Co-investigator AI: The rise of agentic AI for smarter, trustworthy AML compliance narratives, 2025

    Prathamesh Vasudeo Naik, Naresh Kumar Dintakurthi, Zhanghao Hu, Yue Wang, and Robby Qiu. Co-investigator AI: The rise of agentic AI for smarter, trustworthy AML compliance narratives, 2025. arXiv preprint arXiv:2509.08380

  11. [11]

    FAA framework: A large language model-based approach for credit card fraud investigations, 2025

    Shaun Shuster, Eyal Zaloof, Asaf Shabtai, and Rami Puzis. FAA framework: A large language model-based approach for credit card fraud investigations, 2025. arXiv preprint arXiv:2506.11635

  12. [12]

    Agent security bench (ASB): Formalizing and bench- marking attacks and defenses in LLM-based agents, 2025

    Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hong- wei Wang, and Yongfeng Zhang. Agent security bench (ASB): Formalizing and bench- marking attacks and defenses in LLM-based agents, 2025. ICLR 2025; arXiv preprint arXiv:2410.02644

  13. [13]

    PoisonedRAG: Knowledge cor- ruption attacks to retrieval-augmented generation of large language models, 2024

    Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. PoisonedRAG: Knowledge cor- ruption attacks to retrieval-augmented generation of large language models, 2024. USENIX Security 2025; arXiv preprint arXiv:2402.07867

  14. [14]

    Selective conformal risk control, 2025

    Yunpeng Xu, Wenge Guo, and Zhi Wei. Selective conformal risk control, 2025. arXiv preprint arXiv:2512.12844

  15. [15]

    Enhancing the interpretability of SHAP values using LLMs, 2024

    Zeng and Zhu. Enhancing the interpretability of SHAP values using LLMs, 2024. arXiv preprint arXiv:2409.00079

  16. [16]

    Safeguarding large language models: A survey, 2024

    Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, Saddek Bensalem, and Xiaowei Huang. Safeguarding large language models: A survey, 2024. arXiv preprint arXiv:2406.02622; survey on guardrails, input rejection, output moderation, and red-teaming protocols

  17. [17]

    Safety at scale: A comprehensive survey of large model and agent safety, 2025

    Xingjun Ma, Yifeng Gao, Yixu Wang, Ruofan Wang, Xin Wang, Ye Sun, Yifan Ding, Hengyuan Xu, Yunhao Chen, et al. Safety at scale: A comprehensive survey of large model and agent safety, 2025. arXiv preprint arXiv:2502.05206; survey extending safety across vision-language models, diffusion models, and agentic systems

  18. [18]

    A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

    Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

  19. [19]

    Year-over- year developments in financial fraud detection via deep learning, 2025

    Yisong Chen, Chuqing Zhao, Yixin Xu, Chuanhao Nie, and Yixin Zhang. Year-over- year developments in financial fraud detection via deep learning, 2025. arXiv preprint arXiv:2502.00201; survey reviewing deep learning methods for fraud detection 2019–2024

  20. [20]

    Large language models for financial fraud detection: A systematic review of methodologies, performance, and future directions

    Vineet Kumar and Sameer Shaik. Large language models for financial fraud detection: A systematic review of methodologies, performance, and future directions. Zenodo, 2026. Systematic review of 33 LLM financial-fraud studies. 18

  21. [21]

    Applications of AI-based models for online fraud detection and analysis.Crime Science, 14(7), 2025

    Antonis Papasavva, Samantha Lundrigan, Ed Lowther, Shane Johnson, Enrico Mariconti, Anna Markovska, and Nilufer Tuptuk. Applications of AI-based models for online fraud detection and analysis.Crime Science, 14(7), 2025. doi: 10.1186/s40163-025-00248-8. Systematic review of AI and NLP methods for online fraud detection

  22. [22]

    Large language models in the abuse detection pipeline, 2026

    Suraj Kath, Sanket Badhe, Preet Shah, Ashwin Sampathkumar, and Shivani Gupta. Large language models in the abuse detection pipeline, 2026. Lifecycle-oriented review of LLMs across the abuse-detection pipeline

  23. [23]

    FlexGuard: Continuous risk scoring for strictness-adaptive LLM content moderation, 2026

    Zhihao Ding, Jinming Li, Ze Lu, and Jieming Shi. FlexGuard: Continuous risk scoring for strictness-adaptive LLM content moderation, 2026. arXiv preprint arXiv:2602.23636

  24. [24]

    MeasuringwhatLLMsthinktheydo: SHAPfaithfulnessanddeployability on financial tabular classification, 2025

    Saeed AlMarri, Mathieu Ravaut, Kristof Juhasz, Gautier Marti, Hamdan Al Ahbabi, and IbrahimElfadel. MeasuringwhatLLMsthinktheydo: SHAPfaithfulnessanddeployability on financial tabular classification, 2025. arXiv preprint arXiv:2512.00163

  25. [25]

    Guardians and offenders: A survey on harmful content generation and safety mitigation of LLM, 2025

    Chi Zhang, Changjia Zhu, Junjie Xiong, Xiaoran Xu, Lingyao Li, Yao Liu, and Zhuo Lu. Guardians and offenders: A survey on harmful content generation and safety mitigation of LLM, 2025. arXiv preprint arXiv:2508.05775

  26. [26]

    Safety in large reasoning models: A survey, 2025

    Cheng Wang, Yue Liu, Baolong Bi, Duzhen Zhang, Zhong-Zhi Li, Yingwei Ma, Yufei He, Shengju Yu, Xinfeng Li, Junfeng Fang, Jiaheng Zhang, and Bryan Hooi. Safety in large reasoning models: A survey, 2025. arXiv preprint arXiv:2504.17704; EMNLP Findings 2025

  27. [27]

    DGP: A dual-granularity prompting framework for fraud detection with graph-enhanced LLMs, 2025

    Yuan Li, Jun Hu, Bryan Hooi, Bingsheng He, and Cheng Chen. DGP: A dual-granularity prompting framework for fraud detection with graph-enhanced LLMs, 2025. arXiv preprint arXiv:2507.21653

  28. [28]

    Taber, Andreas Damianou, and Mounia Lalmas

    Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri, Henrik Lindström, Daniel R. Taber, Andreas Damianou, and Mounia Lalmas. Policy-as-prompt: Rethinking content moderation in the age of large language models, 2025. ACM FAccT 2025; arXiv preprint arXiv:2502.18695

  29. [29]

    Acomprehensivereviewof LLM-based content moderation: Advancements, challenges, and future directions, 2025

    CongChen, WeiQu, SiSu, YukunFeng, and Tao Li. Acomprehensivereviewof LLM-based content moderation: Advancements, challenges, and future directions, 2025. Knowledge- Based Systems 330:114689

  30. [30]

    Security of LLM-based agents regarding attacks, defenses, and applications: A comprehensive survey, 2025

    Yaxin Tang, Yijia Liu, Jiahe Lan, Zheng Yan, and Erol Gelenbe. Security of LLM-based agents regarding attacks, defenses, and applications: A comprehensive survey, 2025. Infor- mation Fusion (Elsevier)

  31. [31]

    Feng He, Tianqing Zhu, Dayong Ye, Bo Liu, Wanlei Zhou, and Philip S. Yu. The emerged security and privacy of LLM agent: A survey with case studies.ACM Computing Surveys, 58, 2025. arXiv preprint arXiv:2407.19354

  32. [32]

    Content moderation by LLM: From accuracy to legitimacy, 2025

    Tao Huang. Content moderation by LLM: From accuracy to legitimacy, 2025. Artificial Intelligence Review (Springer); arXiv preprint arXiv:2409.03219

  33. [33]

    TELUSdigitaltrustandsafetytrends2025.https://www.telusdigital

    TELUSDigital. TELUSdigitaltrustandsafetytrends2025.https://www.telusdigital. com/insights/trust-and-safety/resource/trust-and-safety-2025, 2025

  34. [34]

    LLMs for explainable AI: A comprehensive survey, 2025

    Ahsan Bilal, David Ebert, and Beiyu Lin. LLMs for explainable AI: A comprehensive survey, 2025. arXiv preprint arXiv:2504.00125. 19

  35. [35]

    Dziemian, M

    M. Dziemian, M. Lin, X. Fu, M. Nowak, N. Winter, E. Jones, A. Zou, L. Ahmad, K. Chaud- huri, et al. How vulnerable are AI agents to indirect prompt injections? insights from a large-scale public competition, 2026. 31-author GraySwan/Anthropic/CMU/Meta study; all 13 tested frontier models vulnerable

  36. [36]

    FRAUDLLM: Zero-shot fraud detection with large language models, 2026

    Mingyang An, Yuan Gao, Jinghan Li, Xiang Wang, and Xiangnan He. FRAUDLLM: Zero-shot fraud detection with large language models, 2026. Springer LNCS

  37. [37]

    Reinforcement learning of large language models for interpretable credit card fraud detection, 2026

    Cooper Lin, Yanting Zhang, Maohao Ran, Wei Xue, Hongwei Fan, Yibo Xu, Zhenglin Wan, Sirui Han, Yike Guo, and Jun Song. Reinforcement learning of large language models for interpretable credit card fraud detection, 2026. arXiv preprint arXiv:2601.05578

  38. [38]

    Telecom fraud detection based on large language models: A multi-role, multi-layer prompting strategy, 2026

    Ding and Zhou. Telecom fraud detection based on large language models: A multi-role, multi-layer prompting strategy, 2026. Applied Sciences 16(1):544; MDPI

  39. [39]

    Telecom fraud recognition based on large language model neuron selection, 2025

    Lanlan Jiang, Cheng Zhang, Xingguo Qin, Ya Zhou, Guanglun Huang, Hui Li, and Jun Li. Telecom fraud recognition based on large language model neuron selection, 2025. Mathe- matics 13(11):1784; MDPI

  40. [40]

    Can LLMs find fraudsters? multi-level LLM enhanced graph fraud detection, 2025

    Tairan Huang, Yili Wang, Qiutong Li, Changlong He, and Jianliang Gao. Can LLMs find fraudsters? multi-level LLM enhanced graph fraud detection, 2025. arXiv preprint arXiv:2507.11997

  41. [41]

    Exploring the in-context learning capabilities of LLMs for money laundering detection in financial graphs, 2025

    Pirmorad. Exploring the in-context learning capabilities of LLMs for money laundering detection in financial graphs, 2025. AI4FCF-ICDM 2025 workshop paper; arXiv preprint arXiv:2507.14785

  42. [42]

    Mirko Franco, Ombretta Gaggi, and Claudio E. Palazzi. Integrating content moderation systems with large language models, 2025. ACM Transactions on the Web 19, 2025

  43. [43]

    Large language models reproduce racial stereotypes when used for text annotation, 2026

    Petter Törnberg. Large language models reproduce racial stereotypes when used for text annotation, 2026. 19-LLM audit; AAVE vs SAE rating gaps mean∼−0.7on profession- alism/toxicity

  44. [44]

    Dialect vs demographics: Quantifying LLM bias from im- plicit linguistic signals vs

    Irti Haq and Belén Saldías. Dialect vs demographics: Quantifying LLM bias from im- plicit linguistic signals vs. explicit user profiles, 2026. Safety/sanitization protections are keyword-dependent and degrade under implicit dialect signals

  45. [45]

    When to invoke: Refining LLM fairness with toxicity assessment, 2026

    Jing Ren, Bowen Li, Ziqi Xu, Renqiang Luo, Shuo Yu, Xin Ye, Haytham Fayek, Xiaodong Li, and Feng Xia. When to invoke: Refining LLM fairness with toxicity assessment, 2026. Inference-time framework flagging demographically sensitive cases for additional assess- ment

  46. [46]

    Longitudinalmonitoring of LLM content moderation of social issues, 2025

    YunlangDai, EmmaLurie, DanaéMetaxa, andSorelleA.Friedler. Longitudinalmonitoring of LLM content moderation of social issues, 2025. AI Watchman longitudinal audit across 400+ social issues; detects unannounced policy shifts

  47. [47]

    Prompt injection at- tacks in large language models and AI agent systems: A comprehensive review of vul- nerabilities, attack vectors, and defense mechanisms, 2026

    Saidakhror Gulyamov, Said Gulyamov, Andrey Rodionov, Rustam Khursanov, Kambarid- din Mekhmonov, Djakhongir Babaev, and Akmaljon Rakhimjonov. Prompt injection at- tacks in large language models and AI agent systems: A comprehensive review of vul- nerabilities, attack vectors, and defense mechanisms, 2026. MDPI Information 17(1):54, 2026

  48. [48]

    Semantic chameleon: Corpus-dependent poisoning attacks and defenses in RAG systems, 2026

    Scott Thornton. Semantic chameleon: Corpus-dependent poisoning attacks and defenses in RAG systems, 2026. Hybrid BM25+vector retrieval reduces gradient-guided poisoning success from 38% to 0%. 20

  49. [49]

    DataexfiltrationfromSlackAIviaindirectpromptinjection

    PromptArmor. DataexfiltrationfromSlackAIviaindirectpromptinjection. PromptArmor disclosure report, 2024

  50. [50]

    ServiceNow Now Assist AI agent vulnerability (BodySnatcher, cve-2025-12420),

    AppOmni. ServiceNow Now Assist AI agent vulnerability (BodySnatcher, cve-2025-12420),

  51. [51]

    Detecting and analyzing prompt abuse in AI tools

    Microsoft Security. Detecting and analyzing prompt abuse in AI tools. Microsoft Security Blog, AI Application Security series, 2026. Enterprise guidance on detecting and mitigating prompt abuse

  52. [52]

    OWASP GenAI exploit round-up report q1 2026.https: //genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/,

    OWASP GenAI Security Project. OWASP GenAI exploit round-up report q1 2026.https: //genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/,

  53. [53]

    Bloomberg News. Hacker used anthropic’s Claude to steal sensitive mexi- can government data.https://www.bloomberg.com/news/articles/2026-02-25/ hacker-used-anthropic-s-claude-to-steal-sensitive-mexican-data, 2026. Indus- try reporting on agentic LLM-as-attacker orchestration; corroborated by SecurityWeek coverage of Monterrey water-utility OT reconnaissance

  54. [54]

    Breaking the chain: A causal analysis of LLM faithfulness to intermediate struc- tures, 2026

    Oleg Somov, Mikhail Chaichuk, Mikhail Seleznyov, Alexander Panchenko, and Elena Tu- tubalina. Breaking the chain: A causal analysis of LLM faithfulness to intermediate struc- tures, 2026. Pearl front-door interventions show CoT structures often do not mediate predictions

  55. [55]

    Multi-layered framework for LLM hallucination mitigation in high-stakes applications: A tutorial, 2025

    Sachin Hiriyanna and Wenbing Zhao. Multi-layered framework for LLM hallucination mitigation in high-stakes applications: A tutorial, 2025. MDPI Computers 14(8):332, 2025

  56. [56]

    GrafanaGhost indirect-injection exfil, Vertex AI Double Agent identity abuse, Flowise RCE

  57. [57]

    Un- equal uncertainty: Rethinking algorithmic interventions for mitigating discrimination from AI, 2025

    Holli Sargeant, Mackenzie Jorgensen, Arina Shah, Adrian Weller, and Umang Bhatt. Un- equal uncertainty: Rethinking algorithmic interventions for mitigating discrimination from AI, 2025. arXiv preprint arXiv:2508.07872

  58. [58]

    Can large language models automate phishing warning explanations? a controlled experiment on effectiveness and user perception, 2025

    Federico Maria Cau, Giuseppe Desolda, Francesco Greco, Lucio Davide Spano, and Luca Viganò. Can large language models automate phishing warning explanations? a controlled experiment on effectiveness and user perception, 2025. Between-subjectsN= 750RCT comparing Claude-3.5/Llama-3.3 feature-based and counterfactual explanations against human-authored warnings

  59. [59]

    R. Yew, S. Xu, S. Saha, F. Fan, K. Ong, X. Wang, K. Sarkar, J. Yang, and J. Guan. Dynamic content moderation in livestreams, 2026. Production-scale MLLM distillation; 67–76% recall at 80% precision; 6–8% reduction in unwanted-stream views in A/B test

  60. [60]

    Learning conformal abstention policies for adaptive risk manage- ment in large language and vision-language models, 2025

    Sina Tayebati, Divake Kumar, Nastaran Darabi, Dinithi Jayasuriya, Ranganath Krishnan, and Amit Ranjan Trivedi. Learning conformal abstention policies for adaptive risk manage- ment in large language and vision-language models, 2025. arXiv preprint arXiv:2502.06884

  61. [61]

    Y. Kang, J. Chen, K. Xu, S. Zhang, L. Guo, R. Pan, J. Revilla, Z. Sun, and Y. Li. Poly-guard: Massive multi-domain safety policy-grounded guardrail dataset, 2026. Eight- domain guardrail benchmark including Finance (fraud, AML) and Cybersecurity (phishing, malware); 19 guardrail models evaluated. 21

  62. [62]

    Domain knowledge-enhanced LLMs for fraud and concept drift detection, 2026

    Ali Şenol, Garima Agrawal, and Huan Liu. Domain knowledge-enhanced LLMs for fraud and concept drift detection, 2026. Electronics 15(3):534; MDPI; arXiv preprint arXiv:2506.21443

  63. [63]

    LLM performance predictors: Learning when to escalate in hybrid human-AI moderation systems, 2026

    Or Bachar, Or Levi, Sardhendu Mishra, Adi Levi, Manpreet Singh Minhas, Justin Miller, Omer Ben-Porat, Eilon Sheetrit, and Jonathan Morra. LLM performance predictors: Learning when to escalate in hybrid human-AI moderation systems, 2026. Meta-model over LLM confidence signals (log-prob, entropy, attribution) for human-review escalation

  64. [64]

    B. Li, Y. Sheng, T. Yang, P. Zhang, and X. Cao. From judgment to interference: Early stopping LLM harmful outputs via streaming content monitoring, 2025. NeurIPS 2025; macro-F1≥0.95 reading only first 18% of generated tokens on average

  65. [65]

    UCCI: Calibrated uncertainty for cost-optimal LLM cascade routing, 2026

    Varun Kotte. UCCI: Calibrated uncertainty for cost-optimal LLM cascade routing, 2026. Isotonic-calibrated cascade routing; 31% inference-cost reduction and ECE 0.12→0.03 on a 75,000-query production NER workload with measured H100 latency

  66. [66]

    Compliance-scored best-of-N guardrail or- chestration for multimodal document generation in payments dispute defense, 2026

    Nataraj Agaram Sundar and Tejas Morabia. Compliance-scored best-of-N guardrail or- chestration for multimodal document generation in payments dispute defense, 2026. Oper- ational readout: 5 attempts within 20s, 91% compliance; +11.0pp dispute win rate from aggregate operational scenario analysis, not a randomized A/B

  67. [67]

    Robust and efficient guardrails with latent reasoning, 2026

    Siddharth Sai, Xiaofei Wen, and Muhao Chen. Robust and efficient guardrails with latent reasoning, 2026. Latent-reasoning guardrail; 12.9×speedup and 22.4×token reduction vs. explicit-reasoning baseline at matched macro-F1

  68. [68]

    Redefining AI red teaming in the agentic era: From weeks to hours, 2026

    Raja Sekhar Rao Dheekonda, Will Pearce, and Nick Landers. Redefining AI red teaming in the agentic era: From weeks to hours, 2026. Agentic red-teaming framework with 45+ at- tacks/450+ transforms/130+ scorers; OWASP/MITRE ATLAS/NIST AI RMF mappings

  69. [69]

    AI agents may always fall for prompt injections,

    Sahar Abdelnabi and Eugene Bagdasarian. AI agents may always fall for prompt injections,

  70. [72]

    SAGE: An LLM- driven self reflective agentic framework for fraud detection, 2026

    Yichen Chen, Siying Li, Yuhang Liang, Lijun Wang, and Renyang Liu. SAGE: An LLM- driven self reflective agentic framework for fraud detection, 2026. Multi-agent fraud detec- tion across payment, e-commerce, and telecom; benchmark-only evaluation on five datasets and five backbones

  71. [74]

    Contextual-integrity reframing of prompt injection; impossibility-style argument against data-instruction separation defenses. 22

  72. [2024]

    arXiv preprint arXiv:2312.02003; High-Confidence Computing 4(2), 2024

  73. [2025]

    Disclosed by AppOmni; patched 30 October 2025

  74. [2026]

    arXiv preprint arXiv:2601.19684