{"id":"59c64045-a038-424a-83ef-89e3d916efa8","arxiv_id":"2504.21574","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of generative AI applications, cyber threats, and regulatory approaches in global finance, with practical recommendations but no new empirical or theoretical contribution.","lead":"This paper surveys how financial institutions use generative AI in customer service, fraud detection, compliance, and investing, and reviews the cyber threats and regulations that accompany it. It is a broad overview for practitioners and policymakers, not a new research result.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's risk narrative rests on unverified secondhand statistics and internally inconsistent citations, so its quantified claims cannot be independently checked.","rationale":"The reader's UNVERDICTED verdict is appropriate for a survey with no novel method or formal claim; the paper's usefulness depends on faithful synthesis. The weakest link is the evidentiary chain for quantified claims. I attempted to find a stronger internal inconsistency: the EU AI Act section (§6.1.2) cites [61] (MAS Project MindForge) and [63] (DORA) for the AI Act itself, and the duplicate numbering [66]-[68] confirms citation hygiene problems. These are not fatal to a practitioner-oriented survey, but they compound the reliance on secondary statistics. An honest stress-test cannot certify the central narrative as verified. I do not think this warrants REJECT—the qualitative observations about GenAI use cases, cyber threats, and regulatory trends align with public knowledge—but it supports keeping the survey UNVERDICTED until the cited evidence is independently checked. Since the reader already reached UNVERDICTED for essentially this reason, my recommendation is UNCHANGED.","tokens_in":27020,"tokens_out":3051,"duration_ms":32341,"concrete_test":"Perform a citation audit: open refs [3], [34], [35], [61], [62], [63], [66], [67], and [68] and check (a) whether each URL resolves, (b) whether the exact statistic or claim appears in the cited source, and (c) whether each reference number maps to exactly one bibliographic entry. Then recompute the strength of the §3.1 risk narrative without any statistic that fails the check; if both the 118% and 26% figures fail, the survey's quantified threat claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central assertion—that GenAI adoption in finance is both high-value and high-risk—is supported largely by industry-sourced statistics that are quoted without primary-source validation. Three load-bearing numbers anchor §1 and §3.1: the $200–340 billion annual banking value (§1, ref [3]), the 118% rise in AI-driven phishing (§3.1, ref [34]), and the 26% executive deepfake-target rate (§3.1, ref [35]). Each is drawn from a secondary press piece or consultant report, and the survey offers no caveat about sampling, definitions, or conflicts of interest. In addition, the reference list is internally inconsistent: [66] is both NIST AI RMF 1.0 (§7a) and an Economic Times story about SBI's LLM (§2.1); [67] is both a stress-testing paper and the Axis Bank AHA page; [68] is both the HDFC Virtual RM page and a MITRE/Microsoft release. This makes the in-text attributions untraceable in places. Because the survey's only quantified evidence for the threat narrative is these figures, the central claim is exactly as strong as those unverified numbers. If the 118% or 26% figures are overstated, the risk section loses its evidentiary basis; if the McKinsey projection is out of scope, the opportunity claim is likewise weakened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey paper reviews the adoption, risks, regulation, and recommended practices for generative AI in financial institutions. It catalogs applications in customer-facing functions, risk and compliance, investment management, developer productivity, and strategic planning; categorizes emerging cybersecurity threats (including AI-generated phishing, deepfakes, and attacks targeting AI systems); outlines a secure AI lifecycle; discusses ethical concerns and governance; reviews regulatory initiatives in the US, EU, India, and Singapore; and concludes with recommendations for practitioners. The central claim is that GenAI offers substantial opportunities in finance while introducing significant cybersecurity and ethical risks that require robust governance and secure adoption practices.","tokens_in":27232,"tokens_out":4157,"duration_ms":41350,"significance":"The paper provides a broad, clearly organized synthesis of a fast-moving area, with useful taxonomy figures, a comprehensive regulatory overview (notably covering India and Singapore), and practical recommendations on secure AI adoption. If the citation and source-quality issues are fixed, it would serve as a reasonable entry point for researchers and practitioners. Its strengths are breadth, clarity, and organization; it does not introduce new methods, data, or verifiable predictions, and its quantified claims are drawn almost entirely from secondary industry sources. The survey would be more valuable if it included critical assessment of the statistics it relies on rather than presenting them as established facts.","major_comments":[{"comment":"The reference list is internally inconsistent, making several in-text attributions untraceable: [66] is used for both the NIST AI RMF in §7a and the Economic Times SBI LLM report in §2.1; [67] is used for both the stress-testing citation in §7b and the Axis Bank AHA page in §2.1; [68] is used for both the MITRE/Microsoft release in §7d and the HDFC Virtual RM page in §2.1. These duplicated numbers must be renumbered and all in-text citations updated before the survey can be evaluated for source support.","section":"§9 References"},{"comment":"The quantitative backbone of the threat narrative consists of three secondary-source statistics quoted without caveats: the 118% rise in AI-driven phishing (ref [34]), the 26% executive deepfake-target rate (ref [35]), and the 90% of companies reporting cyber-fraud in 2024 (also ref [34]). Each is drawn from a trade-press article or consultant release rather than a peer-reviewed or primary data source, and the definitions and samples behind the numbers are not discussed. Since §3.1's conclusion that 'the threat remains substantial' relies on these figures, the authors should either locate primary sources or explicitly frame these as vendor/industry estimates with appropriate hedging.","section":"§3.1"},{"comment":"Claims such as the '20–30% reduction in time-to-market' (ref [26]) and 'recent empirical studies confirm…' (ref [28]) are presented as established findings, but the cited sources are an EY promotional article and a McKinsey survey. The survey should distinguish between vendor/industry claims and academic evidence, and either soften the language or supply the underlying studies.","section":"§2.4 and §2.5"},{"comment":"In §6.1.2, the EU AI Act is cited as [61], but [61] is the MAS Project MindForge page; the EU AI Act reference appears to be [62]. The citation mismatch obscures the regulatory discussion and must be corrected along with the global renumbering.","section":"§6.1.2"}],"minor_comments":[{"comment":"The text 'These treats are further classified' should read 'These threats are further classified.'","section":"§3"},{"comment":"The phrase 'The demonstrate a sample lifecycle in Figure 4' is ungrammatical; it should be 'Figure 4 demonstrates a sample lifecycle.'","section":"§4"},{"comment":"In the conclusion, 'reals the vast potential' should be 'reveals the vast potential.'","section":"§8 Conclusion"},{"comment":"The manuscript repeatedly calls itself 'this chapter' in the abstract and introduction; if it is submitted as a journal article, the wording should be changed to 'this survey' or 'this paper.'","section":"Abstract and §1"},{"comment":"The reference list jumps from [56] to [60]; entries [57]–[59] are missing. A global renumbering pass is needed.","section":"§9 References"},{"comment":"Figure 1 is described as a section overview but no actual figure appears in the provided text; either include it or remove the reference.","section":"Figure 1"},{"comment":"The SEC/CFPB discussion cites [64], which is a robo-advisory trust paper, not a regulatory source; the authors should cite the actual SEC proposed rule or CFPB document.","section":"§6.1.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a broad survey rather than an original technical contribution; the editor may wish to consider whether this fits the journal's expectations for novelty. The repeated 'this chapter' phrasing suggests a book-chapter origin, and the authors should clarify prior publication status. The citation-numbering problems are extensive enough that a careful technical check should be performed before the next round of review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a survey, not a research paper. There is no new method, dataset, or testable claim. What it does well is assemble a recent cross-section of GenAI usage in finance: customer-facing chatbots, fraud detection, compliance, investment, developer productivity, plus a compact summary of the EU AI Act, DORA, RBI FREE-AI, MAS MindForge, and US SEC/CFPB actions. A practitioner entering the field could get a decent orientation from Sections 2 and 6.\n\nThe soft spots are real. The reference list is internally messed up: [66] appears twice (NIST AI RMF and an Economic Times item on SBI's LLM), and the same for [67] and [68]. That makes the in-text attributions in §2.1 for SBI, Axis, and HDFC untraceable unless the reader guesses. This is not fatal to the narrative, but it is exactly the kind of thing a referee should require fixed. Second, the quantified claims anchor the risk narrative on secondary or non-academic sources: the 118% rise in AI-driven phishing [34], the 26% of executives targeted by deepfakes [35], and even the $200–340 billion banking value [3] come from press write-ups of industry reports, and the survey passes them on without caveats on sampling, definitions, or conflicts of interest. Since these numbers are the main evidence for the threats, they should be either upgraded to primary sources or explicitly flagged as vendor/consultant estimates. A minor issue: the 'global' coverage is thinner than the title promises; India gets a lot of space and China is absent.\n\nI don't think the central argument is broken. GenAI in finance is both high-value and high-risk, and the recommendation set (risk frameworks, red teaming, human-in-the-loop, audit trails) is uncontroversial. The paper is honest, and the one self-citation [72] is appropriate.\n\nMy recommendation: if the venue wants surveys, this deserves a serious referee, but the referee should make acceptance conditional on fixing the reference numbering and adding a cautionary note about the non-primary statistics. For a research venue expecting new results, it is a desk reject or a transfer to a surveys track. I would not cite it in my own work until the reference issues are cleaned up.","headline":"A workmanlike but citation-broken survey: no new science, useful as a regulatory and industry map, and only worth peer review if the venue can insist on fixing the reference list and qualifying the industry statistics.","tokens_in":27736,"tokens_out":2863,"would_cite":false,"duration_ms":27914,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative AI in finance brings real gains, but only when institutions add guardrails.","keywords":["Generative AI","Generative AI in Finance","Financial Services","Financial Sector AI Regulation","Adversarial AI Attacks","Financial Fraud and AI","Ethical AI Governance"],"falsifier":"Collect incident-level data from a sample of banks for 2023-2025 and compare the growth in AI-crafted phishing, deepfake fraud attempts, and successful unauthorized transfers against the survey's cited figures; if independent records show flat or much smaller growth, the survey's central claim that generative AI is sharply escalating financial-crime risk would lose support.","tokens_in":26803,"feed_emoji":"🤖","tokens_out":7604,"duration_ms":78613,"temperature":0.7,"pith_summary":"This survey attempts to establish that generative AI is becoming a core operational technology across finance—used in customer-facing chatbots, personalized advisory, fraud detection, compliance summarization, software development, and strategic planning—and that its benefits are real but conditional. The authors argue that the same models that create efficiency also expand the attack surface: AI-generated phishing, deepfake fraud, and malware are rising, while prompt injection, data poisoning, and model extraction target the AI systems themselves. They further argue that regulators are converging on risk-based rules, and that institutions can capture the value only by pairing deployment with governance, validation, red-teaming, explainability, and human oversight.","feed_headline":"Survey: GenAI helps finance only with guardrails","feed_subtitle":"A global review pairs banking's AI use cases with the cyber, bias and regulatory risks they carry.","key_machinery":"The argument is carried by a classification scheme that separates the financial institution's AI adoption surface from its AI risk surface. On the adoption side, the paper groups applications into customer-facing, risk-and-compliance, investment, developer-productivity, and strategic-planning functions and attaches to each the main technical enabler, such as retrieval-augmented generation for advisory or synthetic data for fraud training. On the risk side, it distinguishes threats enabled by generative AI from threats targeting generative AI systems, and maps the latter to an adversarial-tactics catalog covering prompt injection, poisoning, and extraction. The response side is a seven-stage secure AI lifecycle that pairs each threat with a control—red-teaming, model signing, audit trails, and human-in-the-loop thresholds—so that the survey's recommendation to treat AI models like high-risk financial products follows from the classification.","core_discovery":"On the survey's own terms, the discovery is a structured map of a sector mid-transition: generative AI is not a single use case but a horizontal capability already visible in five functional areas, each with distinct technical methods such as retrieval-augmented generation, synthetic data augmentation, long-context summarization, and fine-tuned dialogue models. The paper's central claim is that the benefits are accompanied by a two-sided threat picture: criminals use generative AI to make phishing, impersonation, and malware more effective, while adversarial attackers target the models through prompt injection, data poisoning, model extraction, and adversarial inputs. Because financial decisions are high-stakes and regulated, the survey concludes that adoption is sustainable only through a secure AI lifecycle—secure data collection, validation and red-teaming, access control, monitoring, secure updates, incident response, and continuous governance—plus ethical controls for bias, explainability, privacy, accountability, and human-in-the-loop review.","pith_inferences":["If the survey's direction is right, an implicit consequence is that fraud defense becomes an AI-versus-AI race, with both attackers and defenders generating synthetic content at scale; a testable next step is measuring whether AI-assisted detection systems close the gap opened by AI-crafted phishing.","The survey's regulatory summary suggests a prediction the authors do not spell out: a de facto global standard will form around the strictest regime, so even institutions outside that jurisdiction should prepare for extraterritorial AI compliance.","A concrete way to extend the lifecycle claims is to run controlled red-teaming exercises on a customer-facing financial chatbot and measure how many prompt-injection attempts succeed before and after input sanitization and output filters are added."],"forward_implications":["Banks and fintechs can expect the fastest generative-AI gains in customer service and compliance summarization, where fine-tuned language models and retrieval-augmented generation already show measurable efficiency.","AI-generated phishing and deepfake impersonation will keep making social engineering cheaper and more convincing, so financial institutions should move beyond voice verification and single-factor email checks.","Models trained on internal financial data are at risk of data extraction and poisoning, making model provenance, checksums, and model signing necessary parts of procurement.","Regulators in major jurisdictions are moving toward risk-based AI rules, so institutions that classify their AI uses by risk level and document model cards will be better positioned for compliance.","Treating generative AI as a high-risk model, with pre-deployment validation, monitoring, and human review of critical outputs, is the paper's recommended path to capturing value without uncontrolled losses."],"supporting_citations":[{"why":"Provides the $200-340 billion annual banking value projection that anchors the opportunity narrative.","marker":"[3]"},{"why":"Provides the corporate-strategy survey that frames generative-AI adoption across banking functions.","marker":"[7]"},{"why":"Supplies the 118% rise in AI-driven phishing and deepfake activity that anchors the threat narrative.","marker":"[34]"},{"why":"Supplies the 26% executive deepfake targeting rate used to quantify impersonation risk.","marker":"[35]"},{"why":"Supplies the adversarial-tactics taxonomy that the survey uses to classify prompt injection, poisoning, and extraction attacks.","marker":"[43]"},{"why":"Documents black-hat chatbot services sold on underground forums, grounding the claim that generative AI lowers malware-development barriers.","marker":"[37][38]"},{"why":"Shows that hidden instructions in external content can hijack LLM-integrated applications, grounding the prompt-injection threat.","marker":"[39][40]"},{"why":"Demonstrates that a sabotaged language model can be placed in a public repository, grounding the supply-chain risk claim.","marker":"[42]"}],"fun_headline_variants":["GenAI in finance: survey sees risks, rules, rewards","Finance GenAI: global survey maps threats and controls","Survey: Finance AI needs guardrails to deliver gains","Global survey: rulebook needed for GenAI in finance","GenAI's finance promise hinges on safeguards"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the industry-reported statistics the survey cites—the 118% rise in AI-driven phishing, the $200-340 billion banking value, and the 26% executive deepfake rate—are accurate and representative, because the paper does not independently validate them.","fun_headline_variants_meta":{"raw":{"variants":["GenAI in finance: survey sees risks, rules, rewards","Finance GenAI: global survey maps threats and controls","Survey: Finance AI needs guardrails to deliver gains","Global survey: rulebook needed for GenAI in finance","GenAI's finance promise hinges on safeguards"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001022,"raw_usage":{"total_tokens":4309,"prompt_tokens":943,"completion_tokens":3366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":3289}},"tokens_in":559,"tokens_out":3366,"duration_ms":26354,"temperature":1.0,"reasoning_tokens":3289,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:58:45.947285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect incident-level data from a sample of banks for 2023-2025 and compare the growth in AI-crafted phishing, deepfake fraud attempts, and successful unauthorized transfers against the survey's cited figures; if independent records show flat or much smaller growth, the survey's central claim that generative AI is sharply escalating financial-crime risk would lose support.","supporting_citations":[{"cited_title":"WormGPT” [37] and “FraudGPT","cited_arxiv_id":null,"evidence_quote":"Provides the $200-340 billion annual banking value projection that anchors the opportunity narrative."},{"cited_title":"GenAI and LLM for Financial Institutions: A Corporate Strategic Survey","cited_arxiv_id":null,"evidence_quote":"Provides the corporate-strategy survey that frames generative-AI adoption across banking functions."},{"cited_title":"Generative AI email scams make cyber fraud rampant in 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the 118% rise in AI-driven phishing and deepfake activity that anchors the threat narrative."}],"review_version":1}