{"id":"ef18644f-d848-4e6c-ac7b-a689cb8b72f5","arxiv_id":"2412.00586","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Fully automated AI spear phishing achieved a 54% click-through rate on 101 human participants, matching human experts and far exceeding a 12% control rate.","lead":"Researchers tested whether large language models can run entire spear phishing campaigns, from researching targets to sending emails. The AI system matched human experts, with 54% of recipients clicking a link, compared to 12% for generic scam emails.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The parity claim rests on an unfair benchmark: human experts were limited to one semi-personalized email, so AI's 'matching' 54% may reflect a personalization advantage, not expert-level deception.","rationale":"The experimental core is real and unusually direct: 101 human subjects, randomized groups, live click tracking, and a control arm. The 54% AI versus 12% control gap is large enough that the basic automation capability claim, 'AI-generated personalized phishing outperforms generic phishing,' is not seriously in doubt. However, the paper's headline is specifically a parity claim ('on par with human experts'), and that claim is only as strong as the human-expert benchmark. Section 3.5.2 deliberately gives human experts one semi-personalized email for the whole group, while Section 3.5.3 gives the AI per-target OSINT profiles. The authors acknowledge this asymmetry (Sections 3.4 and 5.1) but present the resulting 54% versus 54% as if it measures relative skill. A human expert allowed to see the same OSINT profile and spend the measured 34 minutes per target could plausibly beat 54%; if so, the most policy-relevant sentence in the abstract and conclusion is unsupported. This is not an accusation of fraud or internal inconsistency; it is an unmeasured confound in the experimental design. The proposed matched human-expert arm would settle it directly. Until then, the parity conclusion should be treated as conditional, which is exactly the reader's verdict; hence the verdict is unchanged.","tokens_in":22957,"tokens_out":7726,"duration_ms":99225,"concrete_test":"Run a matched replication arm in which human phishing experts are given the AI tool's OSINT profiles for each target and a per-target time budget equal to the measured manual process (~34 minutes, Section 5.1.1), and are asked to write one hyper-personalized email per target. Use a fresh sample of comparable participants with the same recruitment and survey flow, the same sender domain, and the same delivery method, then measure click-through. If the human-expert hyper-personalized click rate is statistically indistinguishable from 54%, the original parity claim survives; if it exceeds 54% by more than the original confidence interval, the paper's central 'on par with human experts' conclusion is not supported by the current design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that fully AI-automated phishing is 'on par with human experts' (Abstract; Section 5.1) depends on the human-expert condition being a fair measure of expert phishing ability. It is not controlled for personalization. In Section 3.5.2, the human expert condition used a single semi-personalized email (a generic cross-disciplinary research invitation) sent to all 24 participants, while in Section 3.5.3 the AI condition generated one hyper-personalized email per target from individual OSINT profiles (Section 3.4). The paper explicitly assigns human experts to 'Category 2' (semi-personalized) and the AI tool to 'Category 3' (hyper-personalized), but never tests the counterfactual in which human experts receive the same per-target OSINT and time budget. If a human expert with the same profiles and roughly 34 minutes per target (Section 5.1.1) produces click-through above 54%, the headline parity is an artifact of the constrained benchmark rather than evidence that AI matches human skill. This is the most load-bearing point because the abstract and conclusion promote exactly this parity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a four-arm phishing field experiment with 101 human participants, comparing a control group of generic scam emails, emails written by human phishing experts, fully AI-automated spear phishing emails generated by a custom OSINT-and-email tool, and AI emails with human-in-the-loop intervention. The reported click-through rates are 12%, 54%, 54%, and 56%, respectively. The paper interprets the equal 54% rates for human experts and fully automated AI as evidence that LLMs now perform 'on par with human experts' in spear phishing. It also reports an 88% self-assessed accuracy rate for the AI OSINT profiles, evaluates five LLMs as phishing detectors (with Claude 3.5 Sonnet achieving 97.25% true-positive detection on a 381-email dataset at zero false positives), and presents a stylized economic model estimating that AI-based phishing is more profitable than manual phishing by up to roughly 50 times for large campaigns.","tokens_in":23206,"tokens_out":5085,"duration_ms":52166,"significance":"If the central claims hold, this is a timely and policy-relevant result: it would be one of the first demonstrations that a fully automated LLM pipeline, from OSINT reconnaissance to personalized email generation, can elicit real clicks from human targets at a rate comparable to human experts. The study's strengths include a genuine human-subjects experiment with IRB approval, transparent click-through measurement via per-target tracking links, a comparison with the authors' prior-year results, and an economic model with explicit sensitivity analysis over conversion rates and wage levels. The detection component is also a useful contribution. However, the headline parity claim rests on an asymmetric comparison between hyper-personalized AI emails and a single semi-personalized human-expert email, and the statistical evidence for 'on par' is weaker than the prose suggests. These issues are fixable but require additional analysis or an additional experimental condition.","major_comments":[{"comment":"The central claim that fully automated AI phishing is 'on par with human experts' is confounded by personalization. The human-expert condition used one semi-personalized email (a cross-disciplinary research invitation) sent identically to all 24 participants, whereas the AI condition generated a hyper-personalized email for each target from individual OSINT profiles. The paper itself labels these Category 2 and Category 3 personalization, respectively. Because the experiment never gives human experts the same per-target OSINT information and time budget, the equal 54% click-through rates cannot distinguish 'AI matches human skill' from 'hyper-personalization outperforms semi-personalization.' The abstract and conclusion should either be reframed to state the comparison as AI-hyper-personalized versus human-semi-personalized, or the authors should add a condition in which human experts receive the same per-target profiles and comparable time.","section":"§3.5.2, §3.5.3, §5.1"},{"comment":"No statistical test supports the 'on par' wording. The four click-through rates are point estimates from groups of 24–26 participants; Table 4 reports standard errors around 12 percentage points for the human-expert and AI groups. The 54% versus 54% equality is exactly equal, but the confidence intervals are wide enough to include substantially different underlying rates. The power calculation in §3.2 is designed to detect a large effect, not to establish equivalence. The authors should report confidence intervals for each click-through rate and, if they wish to claim parity, conduct an equivalence or non-inferiority test (e.g., TOST with a pre-specified margin).","section":"§5.1, Table 4"},{"comment":"The headline '88% accurate OSINT' figure is self-assessed by the authors using the subjective categories in Table 2 ('correct and sufficient information'), with no independent raters, no blinding, and no inter-rater reliability measure. Since the claim that the tool performs 'the entire spear phishing process' depends on OSINT quality, this number should be presented as an internal evaluation, not as a validated accuracy measurement. Independent coding of the collected profiles against the participants' actual public information, or at least a second annotator, would strengthen the claim.","section":"§5.1, Table 3; Abstract"},{"comment":"The economic result that AI phishing is up to roughly 50 times more profitable than human phishing depends heavily on the calibrated 'Time spent' value of 1 minute for the fully automated AI group. Unlike the hybrid group's 4:24, which is recorded from actual intervention times (Section 5.1.1), the 1-minute figure appears to be an assumption rather than a measurement; the fully automated pipeline still requires some oversight, and Section 5.1.1 reports that checking each email took about one minute even in the hybrid condition. The profitability ratios and the break-even group sizes in Figure 9 should be accompanied by a sensitivity analysis over plausible per-email time (e.g., 0–5 minutes), since this parameter directly drives the 'up to 50 times' claim.","section":"§6.2, Table 4, Figure 9"}],"minor_comments":[{"comment":"The control-group email was iteratively edited 'to be less suspicious until it was accepted by all tested email clients.' This makes the control a weak baseline for 'ordinary phishing emails,' since it was modified specifically to avoid spam-filter rejection; the authors should acknowledge this as a limitation or describe how representative the final control email is.","section":"§3.5.1, Appendix A.6"},{"comment":"The larger detection dataset includes AI-generated emails produced by the same tool and prompt template that generated the study emails. High detection accuracy on these emails may partly reflect recognition of the tool's own stylistic template rather than general phishing-detection ability. A temporal or source-based holdout would make the detection results more convincing.","section":"§4.2, §5.2"},{"comment":"The full prompt template is withheld 'due to security considerations.' For reproducibility, the authors should provide a redacted or summarized version that preserves the core instruction structure without enabling direct misuse.","section":"§3.6"},{"comment":"There are several typographical and grammatical errors, including 'spar phishing' in Section 8, 'the this' in Section 6, 'Youíll' in Figure 3, and inconsistent hyphenation of 'human-in-the-loop' in places. A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's raw experimental data are valuable and the topic is important, but the parity claim is currently overstated relative to the experimental design. The main fix—either reframing the conclusion or adding a human-expert-hyper-personalized condition—is substantial enough to warrant a major revision rather than a minor one. I would also encourage the editor to ask for the statistical equivalence analysis and independent OSINT evaluation before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the headline result—fully automated LLM spear phishing at 54% click-through, matching human experts—is real as a measured outcome, but the 'on par with human experts' framing is too strong. Human experts were restricted to a single semi-personalized email sent to all 24 targets, while the AI received per-target OSINT profiles and generated a unique hyper-personalized email for each person. That is not a fair comparison of skill. The paper does acknowledge the asymmetry, which I appreciate, but the abstract and conclusion still present it as straightforward parity. That is the main soft spot.\n\nWhat is actually new: an end-to-end tool that handles OSINT scraping, profile construction, email generation, sending, and click tracking with no human in the loop, plus a 101-person human-subjects experiment. That is a meaningful step beyond last year's studies, which reported that AI needed human intervention to match experts. The time-cost measurement—34 minutes of manual effort per target versus roughly a minute and a few cents for the AI—is a useful quantitative contribution. The detection work is also solid: the 'priming for suspicion' finding is simple, reproducible, and practically relevant, and Claude 3.5 Sonnet's 97% true-positive rate with no false positives on the larger set is worth attention.\n\nOther soft spots, in order of importance: the sample is small (24–26 per arm) and self-selected from university recruitment; the 88% OSINT accuracy is scored by the authors themselves with no blinding or inter-rater check; no code or data is released, which limits independent verification; and the economic 'up to 50 times' claim is a calibrated illustration, not a measured result, though the authors are transparent about that. The IRB-approved deception will draw scrutiny, but the debrief and gift-card compensation are defensible.\n\nNet: the empirical core is worth taking seriously. The parity claim needs to be reframed or tested with humans given the same per-target OSINT budget. I would send this to peer review and expect revision, not rejection.","headline":"A genuinely useful empirical data point on fully automated LLM phishing, but the 'on par with human experts' claim is oversold because the human benchmark was deliberately weaker.","tokens_in":23750,"tokens_out":1788,"would_cite":true,"duration_ms":18177,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fully automated AI phishing emails matched human expert phishers in a 101-person trial, with both at 54% click-through.","keywords":["large language models","spear phishing","automated phishing","OSINT reconnaissance","phishing detection","human subjects experiment","cybersecurity economics","social engineering"],"falsifier":"Re-run the experiment with human experts given the same per-target OSINT profiles and enough time to write one personalized email per target; if their click-through rate exceeds the AI-automated rate by more than sampling error, the on-par claim is falsified.","tokens_in":22746,"feed_emoji":"🎣","tokens_out":7338,"duration_ms":150974,"temperature":0.7,"pith_summary":"This paper tests whether large language models can run a spear-phishing campaign end to end—reconnaissance, personalized email writing, sending, and tracking—without human help. In a 101-person experiment, fully automated AI emails achieved a 54% click-through rate, identical to emails written by human experts and far above the 12% control group. The authors argue this is a sharp jump from similar studies a year earlier, when AI needed human-in-the-loop help to match experts. They also report that AI-gathered background profiles were accurate and useful for 88% of targets and that a suspicion-primed language model detected phishing emails with over 97% accuracy and no false positives. If these numbers hold, the marginal cost of a convincing personalized phish has fallen to near zero, which changes who can mount such campaigns and at what scale.","feed_headline":"AI-only phishing matched human experts: 54% clicked in live test","feed_subtitle":"Fully automated emails that scrape public profiles beat a 12% control group and point to cheap, scalable spear phishing.","key_machinery":"The carrying mechanism is a custom AI-agent tool that closes the whole phishing loop. It starts with OSINT reconnaissance: an agent built around GPT-4o searches a target's name, employer, publications, and social media through a search API and text-based browser, and compiles the findings into a vulnerability profile. A prompt-engineering database, currently written by human experts and exceeding 2,000 characters, encodes Cialdini's six influence principles and the V-Triad credibility-and-relevance framework, turning each profile into a personalized email whose sender, deadlines, and references are tailored per target. The tool then sends the emails, embeds a unique tracking URL per recipient, logs clicks, and feeds outcomes back for self-improvement. The key operating distinction is hyper-personalization per target rather than semi-personalization, and the remaining human-in-the-loop cost is small: editing averaged 2:41 minutes per message when needed, with 75% of edits being minor word changes.","core_discovery":"The paper's central claim is that frontier language models have crossed a threshold: a fully automated pipeline can now perform the entire spear-phishing loop at human-expert effectiveness. The load-bearing numbers come from 101 consenting participants split into four groups: control emails clicked at 12%, human-expert emails at 54%, fully AI-automated emails at 54%, and AI emails with minor human editing at 56%. The AI system scraped each target's public digital footprint into a vulnerability profile, used that profile to compose a personalized email with Claude 3.5 Sonnet, and sent it without human intervention; the authors score 88% of the profiles as accurate and useful and only 4% as based on the wrong person. A companion detection experiment finds that Claude 3.5 Sonnet, when explicitly asked whether an email is suspicious, caught 97.25% of 363 phishing emails with zero false positives, outperforming humans on several seemingly benign emails. The paper concludes that AI-automated phishing is now on par with expert human phishing, that guardrail circumvention is trivial, and that the economics favor AI at scale.","pith_inferences":["If the 54% figure transfers beyond a university-recruited sample, the binding constraint on phishing volume is no longer attacker skill but target acquisition and email deliverability; defenses should concentrate on those choke points.","The parity result compares hyper-personalized AI emails against a single semi-personalized human email sent to the whole group; giving human experts the same per-target OSINT profiles would test whether the parity claim is about the AI or about the benchmark.","The OSINT profiles that enable attack could be turned into a personal defensive audit: users could generate their own vulnerability profile, see which public data points a phisher would exploit, and remove data that helps attackers more than it helps them.","A suspicion-primed model with zero false positives suggests an inbox-integrated 'second opinion' feature for suspicious messages, although its usefulness would depend on how often users consult it."],"forward_implications":["Signature-based spam filters lose much of their value because every AI email is unique per target; detection must shift to intent and sender-behavior analysis.","Human-in-the-loop editing is becoming redundant: most edits were minor, and removing them barely changed the click-through rate (54% versus 56%).","The economics of phishing invert: with roughly four cents of API cost per email and a 54% click rate, AI-automated phishing is modeled as profitable for audiences above a few thousand, up to 50 times more profitable than manual phishing.","Suspicion priming is a cheap defensive lever: telling a model to look for suspicious intent rather than asking for the email's intention improved detection without raising false positives.","Because guardrails were bypassed by simple rephrasing, such as changing 'phishing email' to 'email', model providers cannot rely on refusal training alone to stop this use."],"supporting_citations":[{"why":"Supplies the V-Triad credibility-and-relevance framework, send-time best practices, and the free-text response categories used to design and analyze the emails.","marker":"[6]"},{"why":"Supplies the six influence principles (authority, scarcity, liking, and the rest) encoded into the AI prompt templates.","marker":"[48]"},{"why":"Prior-year human-subjects study in which AI phishing needed human-in-the-loop help to match people; the baseline the paper claims to surpass.","marker":"[20]"},{"why":"The authors' earlier LLM phishing study, used as a 2023 baseline showing AI content quality far below this year's results.","marker":"[21]"},{"why":"Provides the 2023 AI-generated email content-quality scores that Table 3 and the performance-growth projection compare against.","marker":"[50]"},{"why":"The GPT-4o model that runs the OSINT reconnaissance agent within the phishing tool.","marker":"[41]"},{"why":"The Claude 3.5 Sonnet model that writes the phishing emails and later performs the intent-detection evaluations.","marker":"[42]"},{"why":"One of the machine-learning phishing-detection benchmarks against which the LLM detectors are compared in the results.","marker":"[51]"},{"why":"Another ML-based phishing detection benchmark used in the comparison plot.","marker":"[52]"},{"why":"A deep-learning phishing-detection survey used as a third baseline in the detection comparison.","marker":"[53]"}],"fun_headline_variants":["AI-only spear phishing hits human expert click rates","Fully automated AI phishing ties human experts at 54%","AI runs full spear phishing pipeline without human help","AI phishing click rate 54%, control group only 12%","AI phishing matches human experts; detection beats humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that AI matches human experts assumes the human-expert arm was a fair test of expert skill, but the experts sent a single semi-personalized email to all 24 targets while the AI wrote a hyper-personalized email for each individual from an OSINT profile.","fun_headline_variants_meta":{"raw":{"variants":["AI-only spear phishing hits human expert click rates","Fully automated AI phishing ties human experts at 54%","AI runs full spear phishing pipeline without human help","AI phishing click rate 54%, control group only 12%","AI phishing matches human experts; detection beats humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001032,"raw_usage":{"total_tokens":4400,"prompt_tokens":1055,"completion_tokens":3345,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":3267}},"tokens_in":671,"tokens_out":3345,"duration_ms":23511,"temperature":1.0,"reasoning_tokens":3267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:12:52.783377+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the experiment with human experts given the same per-target OSINT profiles and enough time to write one personalized email per target; if their click-through rate exceeds the AI-automated rate by more than sampling error, the on-par claim is falsified.","supporting_citations":[{"cited_title":"A comprehensive survey of ai-enabled phishing attacks detection techniques,","cited_arxiv_id":null,"evidence_quote":"Another ML-based phishing detection benchmark used in the comparison plot."},{"cited_title":"Vishwanath, The Weakest Link: How to Diagnose, Detect, and Defend Users from Phishing","cited_arxiv_id":null,"evidence_quote":"Supplies the V-Triad credibility-and-relevance framework, send-time best practices, and the free-text response categories used to design and analyze the emails."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the six influence principles (authority, scarcity, liking, and the rest) encoded into the AI prompt templates."},{"cited_title":"How well does GPT phish people? An investigation involving cognitive biases and feedback,","cited_arxiv_id":null,"evidence_quote":"Prior-year human-subjects study in which AI phishing needed human-in-the-loop help to match people; the baseline the paper claims to surpass."},{"cited_title":"Devising and Detecting Phishing Emails Using Large Language Models,","cited_arxiv_id":null,"evidence_quote":"The authors' earlier LLM phishing study, used as a 2023 baseline showing AI content quality far below this year's results."},{"cited_title":"Gpt-4o: Openai’s language model,","cited_arxiv_id":null,"evidence_quote":"The GPT-4o model that runs the OSINT reconnaissance agent within the phishing tool."},{"cited_title":"Claude 3.5 sonnet: Anthropic’s language model,","cited_arxiv_id":null,"evidence_quote":"The Claude 3.5 Sonnet model that writes the phishing emails and later performs the intent-detection evaluations."},{"cited_title":"Applicability of machine learning in spam and phishing email filtering: review and approaches,","cited_arxiv_id":null,"evidence_quote":"One of the machine-learning phishing-detection benchmarks against which the LLM detectors are compared in the results."},{"cited_title":"Deep learning for phishing detection: Taxonomy, current challenges and future directions,","cited_arxiv_id":null,"evidence_quote":"A deep-learning phishing-detection survey used as a third baseline in the detection comparison."}],"review_version":1}