REVIEW 2 major objections 5 minor 74 references
A survey of 49 sources finds fraud-LLM papers omit latency, cost, and calibration evidence.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:03 UTC pith:S5VGKBAI
load-bearing objection Useful, honestly-scoped evidence audit; the headline fraud-vs-moderation imbalance is real in the coded corpus but partly vulnerable to the author's own search-asymmetry caveat. the 2 major comments →
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is an evidence imbalance in the coded corpus: fraud detection is the largest task-specific category, yet none of its 18 sources report clean per-decision latency, per-decision dollar cost, or calibration. Most fraud work reports offline task performance, retrieval gains, or case-study accuracy; only one source has page-verified batch-runtime/deployment evidence, one reports token usage, and three report human or user-facing evidence. The moderation and abuse sources, though fewer, include livestream A/B results, streaming token-budget results, longitudinal audits, and demographic-fairness audits. The paper argues this imbalance means deployment
What carries the argument
The load-bearing organizing device is FORTE, a role-and-evidence frame that locates an LLM in a workflow by five lenses: fraud-first anchoring, operational placement patterns, roles (classifier, retrieval interface, explanation generator, reviewer assistant, agent, feature extractor, escalation component), trust axes (robustness, explanation integrity, selective prediction), and evaluation-gap analysis. The paper's quantitative engine is a manually coded evidence matrix of 49 sources, with strict definitions separating clean per-decision latency from batch runtime, dollar cost from token counts, and calibration from accuracy. The frame lets the survey compare papers across fraud, moderation,
Load-bearing premise
The survey's headline gap rests on the assumption that the coded corpus fairly represents the public literature; the author explicitly notes that a targeted scan for moderation-side fairness audits was not matched by an equally targeted scan for fraud-side latency and cost reporting, so part of the imbalance could reflect search effort rather than reality.
What would settle it
A targeted search of the same venues and time window using fraud-side latency, cost, and calibration terms—mirroring the moderation-side scan—that retrieves even one fraud or investigation paper reporting clean per-decision latency, per-decision dollar cost, or calibration would falsify the claim that none exist.
If this is right
- If the evidence imbalance is real, public literature does not currently justify deploying an LLM as a real-time fraud scorer without additional latency, cost, and calibration data.
- Selective and hybrid architectures—fast conventional scoring with LLMs reserved for hard cases, explanation, escalation, or investigation—are the most defensible near-term option supported by the corpus.
- Fraud papers that report at least three of latency budget, inference cost, explicit decision threshold, and a human-impact metric would substantially close the gap.
- The minimum deployment-evidence checklist gives practitioners and auditors a concrete standard for judging operational maturity of LLM-based trust-and-safety systems.
- Robustness findings from cross-cutting work, especially indirect prompt injection and retrieval poisoning, should be treated as mandatory evaluation layers for fraud and moderation deployments.
Where Pith is reading between the lines
- The measured fraud–moderation imbalance may partly reflect publication norms: fraud teams in production may withhold latency and cost data for adversarial reasons, so a symmetric search of industry and incident-report channels could surface evidence the academic corpus misses.
- A natural next step is to treat the five-item checklist as a reporting standard for new fraud-LLM papers; if adopted, the field would generate the Pareto-style cost-benefit evidence the survey says is missing.
- The FORTE role vocabulary could transfer to adjacent high-stakes domains such as healthcare triage or legal review, where the same question—what role does the model play and what evidence supports it—applies verbatim.
- Because the paper's counts are frozen at a May 2026 snapshot, its 'no fraud source reports calibration' claim is a moving target; post-window papers already reporting calibration in routing suggest the gap may narrow quickly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a structured narrative survey of 49 operationally relevant sources (18 fraud/investigation, 14 moderation/abuse, 17 cross-cutting robustness) plus 15 contextual references, coding each source for application area, LLM role, adaptation style, deployment pattern, and evidence tags. It introduces FORTE, a role-based organizing frame; identifies four deployment patterns; proposes a five-item minimum deployment-evidence checklist; and lists nine follow-up studies. The headline empirical finding is an evidence imbalance: in the coded corpus, no fraud/investigation source reports clean per-decision latency, per-decision dollar cost, or calibration, while moderation sources contain more explicit operational evidence on latency, cost, governance, and fairness. The paper repeatedly and appropriately limits its claims to the coded corpus and acknowledges major limitations: non-exhaustive search, single-author coding without inter-rater agreement, thin banking-sector coverage, and an asymmetric targeted search in the May 2026 update.
Significance. Because the paper is an evidence audit, its value depends heavily on the representativeness of the corpus and the reliability of the coding. Within those limits, the contribution is useful: the FORTE taxonomy and the minimum deployment-evidence checklist are concrete and practitioner-usable, and the paper avoids the common failure of treating 49 sources as 49 production deployments. The explicit definitions for 'clean latency,' 'dollar cost,' and 'calibration' are a strength, and the claim about the coded corpus is falsifiable if the evidence matrix is made available. The honest limitation section is a credit. However, the central comparative claim (fraud lacks operational evidence, moderation has it) is vulnerable to the acknowledged search asymmetry, and the coding reliability is not independently auditable in the manuscript. These issues are fixable without changing the paper's scope.
major comments (2)
- [§3, Limitations (fourth bullet); §§1 and 6] The absence claim and the fraud–moderation imbalance are the paper's central result, but the corpus-construction asymmetry directly threatens that result. The May 2026 update ran a targeted scan for moderation-side fairness audits and production moderation systems, while no equally targeted scan was run for fraud-side latency/cost reporting. Because the paper draws a comparative conclusion (fraud lacks operational evidence, moderation has it), the measured imbalance could be an artifact of search effort rather than a property of the literature. The paper flags this, but the abstract and Section 6 nevertheless state the imbalance as a finding about the public literature ('moderation papers ... include more explicit public evidence'). To make the central claim load-bearing, either (a) run the symmetric targeted search and report the query families and screening rules, or (b) consistently r
- [§3, 'Coding scheme' and Table 3] The quantitative heart of the paper is the count of evidence types: 0 clean per-decision latency, 0 dollar cost, 0 calibration among 18 fraud sources. These counts depend on single-author judgments about what counts as 'clean,' 'dollar cost,' and 'calibration.' The paper reports no inter-rater reliability, and the coding definitions, while stated, leave room for judgment (e.g., the boundary between 'batch runtime' and 'per-decision latency'; between 'token-usage evidence' and 'dollar cost'). For a survey whose main empirical claim is an absence, the reader should be able to audit the coding. I recommend making the evidence matrix available as supplementary material (currently it is only mentioned as 'included with the source bundle') and having a second coder code a subset with agreement reported. Without this, the zero counts are a claim about the author's coding, not a machine-checkabl
minor comments (5)
- [Abstract and throughout] Typesetting errors with missing spaces appear in the abstract and Section 2 (e.g., 'Weaddressthisquestion', 'mechanisms.Safety', 'TheLLM Security and Privacysurvey'). A careful proofreading pass is needed.
- [§5, Pattern counts] The counts (19 sources in four patterns; 13 other) are not mapped to individual references. Since the placement coding drives the pattern analysis, a supplementary table listing each source and its pattern/role would increase transparency and help readers verify the counts.
- [§6, Evaluation layers] The text says 'One has page-verified batch-runtime/deployment evidence' without naming the source in the paragraph. Please identify the source so readers can verify the claim without consulting the evidence matrix.
- [§7, Signals after the search window] The July 2026 spot check is described, but the query and selection criteria are not provided with the same detail as the May 2026 corpus. Including the same transparency for the spot check would allow readers to assess whether the after-window sources are comparable to the coded corpus.
- [References] Several references are incomplete or inconsistent: [15] lists only 'Zeng and Zhu'; [38] and [41] have partial author listings. Please complete all references.
Circularity Check
No significant circularity: the central claim is an explicitly scoped, externally coded literature measurement, not a derivative of its own inputs.
full rationale
The paper's central claim is an empirical audit of a coded corpus, not a derivation from assumptions. The 18-fraud/14-moderation counts and the absence of clean per-decision latency, dollar cost, or calibration among fraud sources are coding outcomes; the coding categories are defined upfront (Section 6: "Clean latency means per-decision latency, streaming latency, or an explicit latency budget... Dollar cost means a reported monetary cost... Calibration is counted only when a paper evaluates probability calibration, threshold reliability, or coverage-risk behavior") and the claims are explicitly scoped to the corpus: "Among the 18 fraud and investigation sources, none report clean per-decision latency... None report per-decision dollar cost, and none report calibration." No fitted parameter is renamed as a prediction, no equation is constructed from the target result, and there are no self-citations by the author (no Gabani references appear in the bibliography). The acknowledged limitation—the May 2026 update included "a targeted scan for moderation-side fairness audits and production moderation systems, with no equally targeted scan for fraud-side latency/cost reporting"—is a representativeness threat, not circular reasoning; it weakens the generalization from the coded corpus to the literature but does not make the corpus-internal finding tautological. FORTE is an organizing taxonomy and checklist, not an imported uniqueness theorem or an ansatz smuggled in by citation. Because no load-bearing argument reduces to a self-citation, a definitional equivalence, or a fitted input, the honest finding is no circularity, and the score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption The searched corpus (snowball from three neighbor surveys plus targeted queries, 2023–2026) is representative of the LLM fraud/trust-and-safety literature.
- ad hoc to paper The seven-role taxonomy is the right decomposition of LLM placement in workflows.
- domain assumption Deployment evidence is primarily latency, per-decision cost, calibration, threshold, explanation integrity, and adversarial pressure.
- domain assumption Single-author coding without inter-rater reliability is accurate enough for the headline counts.
invented entities (1)
-
FORTE
no independent evidence
read the original abstract
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.
Figures
Reference graph
Works this paper leans on
-
[1]
Understanding structured financial data with LLMs: A case study on fraud detection, 2025
Xuwei Tan, Yao Ma, and Xueru Zhang. Understanding structured financial data with LLMs: A case study on fraud detection, 2025. arXiv preprint arXiv:2512.13040
Pith/arXiv arXiv 2025
-
[2]
FLAG: Fraud detection with LLM-enhanced graph neural network
Chengdong Yang et al. FLAG: Fraud detection with LLM-enhanced graph neural network. InProceedings of KDD, 2025. doi: 10.1145/3711896.3737220
arXiv 2025
-
[3]
BoQu, ZhurongWang, DaisukeYagi, ZhenXu, YangZhao, YinanShan, andFrankZahrad- nik. LLM-enhanced self-evolving reinforcement learning for multi-step e-commerce pay- ment fraud risk detection. InProceedings of ACL (Industry Track), 2025. arXiv preprint arXiv:2509.18719
arXiv 2025
-
[4]
Advanced real-time fraud detection using RAG-based LLMs, 2025
Gurjot Singh, Prabhjot Singh, and Maninder Singh. Advanced real-time fraud detection using RAG-based LLMs, 2025. arXiv preprint arXiv:2501.15290
Pith/arXiv arXiv 2025
-
[5]
LLM-assistedauthenticationandfrauddetection,
EmunahS-S.ChanandAldarC-F.Chan. LLM-assistedauthenticationandfrauddetection,
-
[6]
AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts, 2024
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien. AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts, 2024. arXiv preprint arXiv:2404.05993. 17
Pith/arXiv arXiv 2024
-
[7]
SLM-Mod: Small language models surpass LLMs at content moderation
Xianyang Zhan, Agam Goyal, Yilun Chen, Eshwar Chandrasekharan, and Koustuv Saha. SLM-Mod: Small language models surpass LLMs at content moderation. InProceedings of NAACL, 2025. arXiv preprint arXiv:2410.13155
Pith/arXiv arXiv 2025
-
[8]
Agentic AI framework for enhancing scam intelligence in digital payments, 2025
Nitish Jaipuria, Lorenzo Gatto, Zijun Kan, Shankey Poddar, Bill Cheung, Diksha Bansal, Ramanan Balakrishnan, Aviral Suri, and Jose Estevez. Agentic AI framework for enhancing scam intelligence in digital payments, 2025. arXiv preprint arXiv:2508.19932
Pith/arXiv arXiv 2025
-
[9]
Bryan Lim, Roman Huerta, Alejandro Sotelo, Anthonie Quintela, and Priyanka Kumar. EXPLICATE: Enhancing phishing detection through explainable AI and LLM-powered interpretability, 2025. arXiv preprint arXiv:2503.20796
Pith/arXiv arXiv 2025
-
[10]
Co-investigator AI: The rise of agentic AI for smarter, trustworthy AML compliance narratives, 2025
Prathamesh Vasudeo Naik, Naresh Kumar Dintakurthi, Zhanghao Hu, Yue Wang, and Robby Qiu. Co-investigator AI: The rise of agentic AI for smarter, trustworthy AML compliance narratives, 2025. arXiv preprint arXiv:2509.08380
arXiv 2025
-
[11]
FAA framework: A large language model-based approach for credit card fraud investigations, 2025
Shaun Shuster, Eyal Zaloof, Asaf Shabtai, and Rami Puzis. FAA framework: A large language model-based approach for credit card fraud investigations, 2025. arXiv preprint arXiv:2506.11635
Pith/arXiv arXiv 2025
-
[12]
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hong- wei Wang, and Yongfeng Zhang. Agent security bench (ASB): Formalizing and bench- marking attacks and defenses in LLM-based agents, 2025. ICLR 2025; arXiv preprint arXiv:2410.02644
Pith/arXiv arXiv 2025
-
[13]
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. PoisonedRAG: Knowledge cor- ruption attacks to retrieval-augmented generation of large language models, 2024. USENIX Security 2025; arXiv preprint arXiv:2402.07867
Pith/arXiv arXiv 2024
-
[14]
Selective conformal risk control, 2025
Yunpeng Xu, Wenge Guo, and Zhi Wei. Selective conformal risk control, 2025. arXiv preprint arXiv:2512.12844
Pith/arXiv arXiv 2025
-
[15]
Enhancing the interpretability of SHAP values using LLMs, 2024
Zeng and Zhu. Enhancing the interpretability of SHAP values using LLMs, 2024. arXiv preprint arXiv:2409.00079
arXiv 2024
-
[16]
Safeguarding large language models: A survey, 2024
Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, Saddek Bensalem, and Xiaowei Huang. Safeguarding large language models: A survey, 2024. arXiv preprint arXiv:2406.02622; survey on guardrails, input rejection, output moderation, and red-teaming protocols
Pith/arXiv arXiv 2024
-
[17]
Safety at scale: A comprehensive survey of large model and agent safety, 2025
Xingjun Ma, Yifeng Gao, Yixu Wang, Ruofan Wang, Xin Wang, Ye Sun, Yifan Ding, Hengyuan Xu, Yunhao Chen, et al. Safety at scale: A comprehensive survey of large model and agent safety, 2025. arXiv preprint arXiv:2502.05206; survey extending safety across vision-language models, diffusion models, and agentic systems
Pith/arXiv arXiv 2025
-
[18]
A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,
-
[19]
Year-over- year developments in financial fraud detection via deep learning, 2025
Yisong Chen, Chuqing Zhao, Yixin Xu, Chuanhao Nie, and Yixin Zhang. Year-over- year developments in financial fraud detection via deep learning, 2025. arXiv preprint arXiv:2502.00201; survey reviewing deep learning methods for fraud detection 2019–2024
Pith/arXiv arXiv 2025
-
[20]
Large language models for financial fraud detection: A systematic review of methodologies, performance, and future directions
Vineet Kumar and Sameer Shaik. Large language models for financial fraud detection: A systematic review of methodologies, performance, and future directions. Zenodo, 2026. Systematic review of 33 LLM financial-fraud studies. 18
2026
-
[21]
Applications of AI-based models for online fraud detection and analysis.Crime Science, 14(7), 2025
Antonis Papasavva, Samantha Lundrigan, Ed Lowther, Shane Johnson, Enrico Mariconti, Anna Markovska, and Nilufer Tuptuk. Applications of AI-based models for online fraud detection and analysis.Crime Science, 14(7), 2025. doi: 10.1186/s40163-025-00248-8. Systematic review of AI and NLP methods for online fraud detection
-
[22]
Large language models in the abuse detection pipeline, 2026
Suraj Kath, Sanket Badhe, Preet Shah, Ashwin Sampathkumar, and Shivani Gupta. Large language models in the abuse detection pipeline, 2026. Lifecycle-oriented review of LLMs across the abuse-detection pipeline
2026
-
[23]
FlexGuard: Continuous risk scoring for strictness-adaptive LLM content moderation, 2026
Zhihao Ding, Jinming Li, Ze Lu, and Jieming Shi. FlexGuard: Continuous risk scoring for strictness-adaptive LLM content moderation, 2026. arXiv preprint arXiv:2602.23636
Pith/arXiv arXiv 2026
-
[24]
Saeed AlMarri, Mathieu Ravaut, Kristof Juhasz, Gautier Marti, Hamdan Al Ahbabi, and IbrahimElfadel. MeasuringwhatLLMsthinktheydo: SHAPfaithfulnessanddeployability on financial tabular classification, 2025. arXiv preprint arXiv:2512.00163
arXiv 2025
-
[25]
Guardians and offenders: A survey on harmful content generation and safety mitigation of LLM, 2025
Chi Zhang, Changjia Zhu, Junjie Xiong, Xiaoran Xu, Lingyao Li, Yao Liu, and Zhuo Lu. Guardians and offenders: A survey on harmful content generation and safety mitigation of LLM, 2025. arXiv preprint arXiv:2508.05775
Pith/arXiv arXiv 2025
-
[26]
Safety in large reasoning models: A survey, 2025
Cheng Wang, Yue Liu, Baolong Bi, Duzhen Zhang, Zhong-Zhi Li, Yingwei Ma, Yufei He, Shengju Yu, Xinfeng Li, Junfeng Fang, Jiaheng Zhang, and Bryan Hooi. Safety in large reasoning models: A survey, 2025. arXiv preprint arXiv:2504.17704; EMNLP Findings 2025
Pith/arXiv arXiv 2025
-
[27]
DGP: A dual-granularity prompting framework for fraud detection with graph-enhanced LLMs, 2025
Yuan Li, Jun Hu, Bryan Hooi, Bingsheng He, and Cheng Chen. DGP: A dual-granularity prompting framework for fraud detection with graph-enhanced LLMs, 2025. arXiv preprint arXiv:2507.21653
Pith/arXiv arXiv 2025
-
[28]
Taber, Andreas Damianou, and Mounia Lalmas
Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri, Henrik Lindström, Daniel R. Taber, Andreas Damianou, and Mounia Lalmas. Policy-as-prompt: Rethinking content moderation in the age of large language models, 2025. ACM FAccT 2025; arXiv preprint arXiv:2502.18695
Pith/arXiv arXiv 2025
-
[29]
Acomprehensivereviewof LLM-based content moderation: Advancements, challenges, and future directions, 2025
CongChen, WeiQu, SiSu, YukunFeng, and Tao Li. Acomprehensivereviewof LLM-based content moderation: Advancements, challenges, and future directions, 2025. Knowledge- Based Systems 330:114689
2025
-
[30]
Security of LLM-based agents regarding attacks, defenses, and applications: A comprehensive survey, 2025
Yaxin Tang, Yijia Liu, Jiahe Lan, Zheng Yan, and Erol Gelenbe. Security of LLM-based agents regarding attacks, defenses, and applications: A comprehensive survey, 2025. Infor- mation Fusion (Elsevier)
2025
-
[31]
Feng He, Tianqing Zhu, Dayong Ye, Bo Liu, Wanlei Zhou, and Philip S. Yu. The emerged security and privacy of LLM agent: A survey with case studies.ACM Computing Surveys, 58, 2025. arXiv preprint arXiv:2407.19354
arXiv 2025
-
[32]
Content moderation by LLM: From accuracy to legitimacy, 2025
Tao Huang. Content moderation by LLM: From accuracy to legitimacy, 2025. Artificial Intelligence Review (Springer); arXiv preprint arXiv:2409.03219
Pith/arXiv arXiv 2025
-
[33]
TELUSdigitaltrustandsafetytrends2025.https://www.telusdigital
TELUSDigital. TELUSdigitaltrustandsafetytrends2025.https://www.telusdigital. com/insights/trust-and-safety/resource/trust-and-safety-2025, 2025
2025
-
[34]
LLMs for explainable AI: A comprehensive survey, 2025
Ahsan Bilal, David Ebert, and Beiyu Lin. LLMs for explainable AI: A comprehensive survey, 2025. arXiv preprint arXiv:2504.00125. 19
Pith/arXiv arXiv 2025
-
[35]
Dziemian, M
M. Dziemian, M. Lin, X. Fu, M. Nowak, N. Winter, E. Jones, A. Zou, L. Ahmad, K. Chaud- huri, et al. How vulnerable are AI agents to indirect prompt injections? insights from a large-scale public competition, 2026. 31-author GraySwan/Anthropic/CMU/Meta study; all 13 tested frontier models vulnerable
2026
-
[36]
FRAUDLLM: Zero-shot fraud detection with large language models, 2026
Mingyang An, Yuan Gao, Jinghan Li, Xiang Wang, and Xiangnan He. FRAUDLLM: Zero-shot fraud detection with large language models, 2026. Springer LNCS
2026
-
[37]
Reinforcement learning of large language models for interpretable credit card fraud detection, 2026
Cooper Lin, Yanting Zhang, Maohao Ran, Wei Xue, Hongwei Fan, Yibo Xu, Zhenglin Wan, Sirui Han, Yike Guo, and Jun Song. Reinforcement learning of large language models for interpretable credit card fraud detection, 2026. arXiv preprint arXiv:2601.05578
arXiv 2026
-
[38]
Telecom fraud detection based on large language models: A multi-role, multi-layer prompting strategy, 2026
Ding and Zhou. Telecom fraud detection based on large language models: A multi-role, multi-layer prompting strategy, 2026. Applied Sciences 16(1):544; MDPI
2026
-
[39]
Telecom fraud recognition based on large language model neuron selection, 2025
Lanlan Jiang, Cheng Zhang, Xingguo Qin, Ya Zhou, Guanglun Huang, Hui Li, and Jun Li. Telecom fraud recognition based on large language model neuron selection, 2025. Mathe- matics 13(11):1784; MDPI
2025
-
[40]
Can LLMs find fraudsters? multi-level LLM enhanced graph fraud detection, 2025
Tairan Huang, Yili Wang, Qiutong Li, Changlong He, and Jianliang Gao. Can LLMs find fraudsters? multi-level LLM enhanced graph fraud detection, 2025. arXiv preprint arXiv:2507.11997
arXiv 2025
-
[41]
Pirmorad. Exploring the in-context learning capabilities of LLMs for money laundering detection in financial graphs, 2025. AI4FCF-ICDM 2025 workshop paper; arXiv preprint arXiv:2507.14785
Pith/arXiv arXiv 2025
-
[42]
Mirko Franco, Ombretta Gaggi, and Claudio E. Palazzi. Integrating content moderation systems with large language models, 2025. ACM Transactions on the Web 19, 2025
2025
-
[43]
Large language models reproduce racial stereotypes when used for text annotation, 2026
Petter Törnberg. Large language models reproduce racial stereotypes when used for text annotation, 2026. 19-LLM audit; AAVE vs SAE rating gaps mean∼−0.7on profession- alism/toxicity
2026
-
[44]
Dialect vs demographics: Quantifying LLM bias from im- plicit linguistic signals vs
Irti Haq and Belén Saldías. Dialect vs demographics: Quantifying LLM bias from im- plicit linguistic signals vs. explicit user profiles, 2026. Safety/sanitization protections are keyword-dependent and degrade under implicit dialect signals
2026
-
[45]
When to invoke: Refining LLM fairness with toxicity assessment, 2026
Jing Ren, Bowen Li, Ziqi Xu, Renqiang Luo, Shuo Yu, Xin Ye, Haytham Fayek, Xiaodong Li, and Feng Xia. When to invoke: Refining LLM fairness with toxicity assessment, 2026. Inference-time framework flagging demographically sensitive cases for additional assess- ment
2026
-
[46]
Longitudinalmonitoring of LLM content moderation of social issues, 2025
YunlangDai, EmmaLurie, DanaéMetaxa, andSorelleA.Friedler. Longitudinalmonitoring of LLM content moderation of social issues, 2025. AI Watchman longitudinal audit across 400+ social issues; detects unannounced policy shifts
2025
-
[47]
Prompt injection at- tacks in large language models and AI agent systems: A comprehensive review of vul- nerabilities, attack vectors, and defense mechanisms, 2026
Saidakhror Gulyamov, Said Gulyamov, Andrey Rodionov, Rustam Khursanov, Kambarid- din Mekhmonov, Djakhongir Babaev, and Akmaljon Rakhimjonov. Prompt injection at- tacks in large language models and AI agent systems: A comprehensive review of vul- nerabilities, attack vectors, and defense mechanisms, 2026. MDPI Information 17(1):54, 2026
2026
-
[48]
Semantic chameleon: Corpus-dependent poisoning attacks and defenses in RAG systems, 2026
Scott Thornton. Semantic chameleon: Corpus-dependent poisoning attacks and defenses in RAG systems, 2026. Hybrid BM25+vector retrieval reduces gradient-guided poisoning success from 38% to 0%. 20
2026
-
[49]
DataexfiltrationfromSlackAIviaindirectpromptinjection
PromptArmor. DataexfiltrationfromSlackAIviaindirectpromptinjection. PromptArmor disclosure report, 2024
2024
-
[50]
ServiceNow Now Assist AI agent vulnerability (BodySnatcher, cve-2025-12420),
AppOmni. ServiceNow Now Assist AI agent vulnerability (BodySnatcher, cve-2025-12420),
2025
-
[51]
Detecting and analyzing prompt abuse in AI tools
Microsoft Security. Detecting and analyzing prompt abuse in AI tools. Microsoft Security Blog, AI Application Security series, 2026. Enterprise guidance on detecting and mitigating prompt abuse
2026
-
[52]
OWASP GenAI exploit round-up report q1 2026.https: //genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/,
OWASP GenAI Security Project. OWASP GenAI exploit round-up report q1 2026.https: //genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/,
2026
-
[53]
Bloomberg News. Hacker used anthropic’s Claude to steal sensitive mexi- can government data.https://www.bloomberg.com/news/articles/2026-02-25/ hacker-used-anthropic-s-claude-to-steal-sensitive-mexican-data, 2026. Indus- try reporting on agentic LLM-as-attacker orchestration; corroborated by SecurityWeek coverage of Monterrey water-utility OT reconnaissance
2026
-
[54]
Breaking the chain: A causal analysis of LLM faithfulness to intermediate struc- tures, 2026
Oleg Somov, Mikhail Chaichuk, Mikhail Seleznyov, Alexander Panchenko, and Elena Tu- tubalina. Breaking the chain: A causal analysis of LLM faithfulness to intermediate struc- tures, 2026. Pearl front-door interventions show CoT structures often do not mediate predictions
2026
-
[55]
Multi-layered framework for LLM hallucination mitigation in high-stakes applications: A tutorial, 2025
Sachin Hiriyanna and Wenbing Zhao. Multi-layered framework for LLM hallucination mitigation in high-stakes applications: A tutorial, 2025. MDPI Computers 14(8):332, 2025
2025
-
[56]
GrafanaGhost indirect-injection exfil, Vertex AI Double Agent identity abuse, Flowise RCE
-
[57]
Holli Sargeant, Mackenzie Jorgensen, Arina Shah, Adrian Weller, and Umang Bhatt. Un- equal uncertainty: Rethinking algorithmic interventions for mitigating discrimination from AI, 2025. arXiv preprint arXiv:2508.07872
Pith/arXiv arXiv 2025
-
[58]
Can large language models automate phishing warning explanations? a controlled experiment on effectiveness and user perception, 2025
Federico Maria Cau, Giuseppe Desolda, Francesco Greco, Lucio Davide Spano, and Luca Viganò. Can large language models automate phishing warning explanations? a controlled experiment on effectiveness and user perception, 2025. Between-subjectsN= 750RCT comparing Claude-3.5/Llama-3.3 feature-based and counterfactual explanations against human-authored warnings
2025
-
[59]
R. Yew, S. Xu, S. Saha, F. Fan, K. Ong, X. Wang, K. Sarkar, J. Yang, and J. Guan. Dynamic content moderation in livestreams, 2026. Production-scale MLLM distillation; 67–76% recall at 80% precision; 6–8% reduction in unwanted-stream views in A/B test
2026
-
[60]
Sina Tayebati, Divake Kumar, Nastaran Darabi, Dinithi Jayasuriya, Ranganath Krishnan, and Amit Ranjan Trivedi. Learning conformal abstention policies for adaptive risk manage- ment in large language and vision-language models, 2025. arXiv preprint arXiv:2502.06884
Pith/arXiv arXiv 2025
-
[61]
Y. Kang, J. Chen, K. Xu, S. Zhang, L. Guo, R. Pan, J. Revilla, Z. Sun, and Y. Li. Poly-guard: Massive multi-domain safety policy-grounded guardrail dataset, 2026. Eight- domain guardrail benchmark including Finance (fraud, AML) and Cybersecurity (phishing, malware); 19 guardrail models evaluated. 21
2026
-
[62]
Domain knowledge-enhanced LLMs for fraud and concept drift detection, 2026
Ali Şenol, Garima Agrawal, and Huan Liu. Domain knowledge-enhanced LLMs for fraud and concept drift detection, 2026. Electronics 15(3):534; MDPI; arXiv preprint arXiv:2506.21443
Pith/arXiv arXiv 2026
-
[63]
LLM performance predictors: Learning when to escalate in hybrid human-AI moderation systems, 2026
Or Bachar, Or Levi, Sardhendu Mishra, Adi Levi, Manpreet Singh Minhas, Justin Miller, Omer Ben-Porat, Eilon Sheetrit, and Jonathan Morra. LLM performance predictors: Learning when to escalate in hybrid human-AI moderation systems, 2026. Meta-model over LLM confidence signals (log-prob, entropy, attribution) for human-review escalation
2026
-
[64]
B. Li, Y. Sheng, T. Yang, P. Zhang, and X. Cao. From judgment to interference: Early stopping LLM harmful outputs via streaming content monitoring, 2025. NeurIPS 2025; macro-F1≥0.95 reading only first 18% of generated tokens on average
2025
-
[65]
UCCI: Calibrated uncertainty for cost-optimal LLM cascade routing, 2026
Varun Kotte. UCCI: Calibrated uncertainty for cost-optimal LLM cascade routing, 2026. Isotonic-calibrated cascade routing; 31% inference-cost reduction and ECE 0.12→0.03 on a 75,000-query production NER workload with measured H100 latency
2026
-
[66]
Compliance-scored best-of-N guardrail or- chestration for multimodal document generation in payments dispute defense, 2026
Nataraj Agaram Sundar and Tejas Morabia. Compliance-scored best-of-N guardrail or- chestration for multimodal document generation in payments dispute defense, 2026. Oper- ational readout: 5 attempts within 20s, 91% compliance; +11.0pp dispute win rate from aggregate operational scenario analysis, not a randomized A/B
2026
-
[67]
Robust and efficient guardrails with latent reasoning, 2026
Siddharth Sai, Xiaofei Wen, and Muhao Chen. Robust and efficient guardrails with latent reasoning, 2026. Latent-reasoning guardrail; 12.9×speedup and 22.4×token reduction vs. explicit-reasoning baseline at matched macro-F1
2026
-
[68]
Redefining AI red teaming in the agentic era: From weeks to hours, 2026
Raja Sekhar Rao Dheekonda, Will Pearce, and Nick Landers. Redefining AI red teaming in the agentic era: From weeks to hours, 2026. Agentic red-teaming framework with 45+ at- tacks/450+ transforms/130+ scorers; OWASP/MITRE ATLAS/NIST AI RMF mappings
2026
-
[69]
AI agents may always fall for prompt injections,
Sahar Abdelnabi and Eugene Bagdasarian. AI agents may always fall for prompt injections,
-
[72]
SAGE: An LLM- driven self reflective agentic framework for fraud detection, 2026
Yichen Chen, Siying Li, Yuhang Liang, Lijun Wang, and Renyang Liu. SAGE: An LLM- driven self reflective agentic framework for fraud detection, 2026. Multi-agent fraud detec- tion across payment, e-commerce, and telecom; benchmark-only evaluation on five datasets and five backbones
2026
-
[74]
Contextual-integrity reframing of prompt injection; impossibility-style argument against data-instruction separation defenses. 22
-
[2024]
arXiv preprint arXiv:2312.02003; High-Confidence Computing 4(2), 2024
Pith/arXiv arXiv 2024
-
[2025]
Disclosed by AppOmni; patched 30 October 2025
2025
-
[2026]
arXiv preprint arXiv:2601.19684
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.